SlideAgent: New AI Framework Improves Analysis of Complex Visual Documents
- Morgan have developed an artificial intelligence framework named SlideAgent to help large language models interpret complex visual documents such as presentation slide decks, brochures, and reports, according to...
- A network of agents, each specialized for a specific level of the document, divides and coordinates the analysis, according to Georgia Tech reporting.
- During their evaluations, the team assessed the novel framework across diverse practical materials, encompassing technical slides, visual question-answering datasets, and financial presentations.
Researchers at Georgia Tech and J.P. Morgan have developed an artificial intelligence framework named SlideAgent to help large language models interpret complex visual documents such as presentation slide decks, brochures, and reports, according to source material from Georgia Tech.
Modern AI platforms can summarize reports and analyze documents quickly, but they frequently miss critical details when information is scattered across dozens of slides, charts, and tables. Yiqiao (Ahren) Jin, a Ph.D. candidate in Georgia Tech’s School of Computational Science and Engineering (CSE), notes that in high-stakes fields like finance, misreading a number, overlooking a footnote, or making an incorrect comparison across pages can affect reporting, risk assessment, and strategic decisions.
Current multimodal artificial intelligence systems often process entire pages at single units. This approach leads to errors such as miscounting items in a chart or overlooking details in dense visuals. SlideAgent addresses this limitation by mimicking how humans naturally read long presentations, breaking documents into three distinct tiers: the full document, individual pages, and specific elements like charts, tables, and text blocks.
How the SlideAgent Framework Operates
A network of agents, each specialized for a specific level of the document, divides and coordinates the analysis, according to Georgia Tech reporting. SlideAgent then combines these outputs to build a structured understanding of the overall document, enabling more accurate answers and reasoning across multiple pages.
“The central inspiration was how people naturally read a long presentation,” Jin said, as reported by Georgia Tech. “We first develop an understanding of the overall narrative, then identify the relevant pages or sections, and finally zoom in on individual charts, tables, or text blocks when precise evidence is needed.”
Evaluation Results and Performance Gains
During their evaluations, the team assessed the novel framework across diverse practical materials, encompassing technical slides, visual question-answering datasets, and financial presentations. SlideAgent consistently outperformed leading commercial models and open-source tools during the evaluations, improving accuracy by up to 10% on complex tasks like cross-slide comparisons.
“We found the results very encouraging,” Jin said in the Georgia Tech report. “SlideAgent reaches an improvement of 7.9% over its proprietary base model and 9.8% over the evaluated open-source base models. These are meaningful gains given the strength of the underlying multimodal models and the difficulty of the tasks.”
The Association for Computational Linguistics has accepted SlideAgent for presentation at its annual meeting, highlighting a broader shift in artificial intelligence research toward smarter design and more efficient reasoning rather than simply building larger models.
