SlideAgent improves AI reading of complex documents in evaluation
A footnote disappears in the corner of a crowded slide. A number is counted incorrectly in a chart. For an AI reading a financial presentation, those small misses can affect reporting, risk assessment or strategic decisions. Researchers from Georgia Tech and J.P. Morgan say their new framework, SlideAgent, improved accuracy by up to 10% on complex visual documents.
The idea is to make AI read more like a person. Current multimodal large language models—systems that work with both text and images—often process an entire page at once. SlideAgent instead asks separate agents to study three levels: the full document, individual pages, and elements such as charts, tables and text blocks. Their findings are then combined into a structured understanding that can support comparisons across pages.
The team tested the framework on real-world financial presentations, technical slides and visual question-answering datasets. SlideAgent delivered a 7.9% improvement over its proprietary base model and a 9.8% improvement over the evaluated open-source base models, according to lead researcher Yiqiao (Ahren) Jin. The strongest gains came on tasks that required comparing information across slides or connecting visuals on the same page.
And then? For people working with dense reports, the practical promise is narrower but useful: fewer missed details when an AI summarizes material or answers questions about it. That could reduce the mental effort involved in navigating long presentations, while giving users a better basis for checking high-stakes information. The work does not establish that SlideAgent is already deployed in offices; it reports results from an evaluation.
The project was accepted for presentation at the Association for Computational Linguistics’ annual meeting, ACL 2026, held in San Diego from July 2–7. Jin, a Georgia Tech Ph.D. candidate, developed SlideAgent during an internship at J.P. Morgan AI Research. The team’s next challenge is explicit: improving computational efficiency without sacrificing accuracy or interpretability.
Comentários
A carregar a conversa…
Inicie sessão para escrever um comentário. Iniciar sessão