SlideAgent: Hierarchical Agentic Framework for Multi-Page Visual Document Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Jin, Yiqiao, Kaur, Rachneet, Zeng, Zhen, Ganesh, Sumitra, Kumar, Srijan |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering
by: Kaur, Rachneet, et al.
Published: (2025)
by: Kaur, Rachneet, et al.
Published: (2025)
AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations
by: Verma, Gaurav, et al.
Published: (2024)
by: Verma, Gaurav, et al.
Published: (2024)
Evaluating Large Language Models on Time Series Feature Understanding: A Comprehensive Taxonomy and Benchmark
by: Fons, Elizabeth, et al.
Published: (2024)
by: Fons, Elizabeth, et al.
Published: (2024)
LETS-C: Leveraging Text Embedding for Time Series Classification
by: Kaur, Rachneet, et al.
Published: (2024)
by: Kaur, Rachneet, et al.
Published: (2024)
Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
by: Sharma, Kartik, et al.
Published: (2025)
by: Sharma, Kartik, et al.
Published: (2025)
AI Analyst: Framework and Comprehensive Evaluation of Large Language Models for Financial Time Series Report Generation
by: Fons, Elizabeth, et al.
Published: (2025)
by: Fons, Elizabeth, et al.
Published: (2025)
MM-Soc: Benchmarking Multimodal Large Language Models in Social Media Platforms
by: Jin, Yiqiao, et al.
Published: (2024)
by: Jin, Yiqiao, et al.
Published: (2024)
LAW: Legal Agentic Workflows for Custody and Fund Services Contracts
by: Watson, William, et al.
Published: (2024)
by: Watson, William, et al.
Published: (2024)
TADACap: Time-series Adaptive Domain-Aware Captioning
by: Fons, Elizabeth, et al.
Published: (2025)
by: Fons, Elizabeth, et al.
Published: (2025)
MedHalu: Hallucinations in Responses to Healthcare Queries by Large Language Models
by: Agarwal, Vibhor, et al.
Published: (2024)
by: Agarwal, Vibhor, et al.
Published: (2024)
DocLens : A Tool-Augmented Multi-Agent Framework for Long Visual Document Understanding
by: Zhu, Dawei, et al.
Published: (2025)
by: Zhu, Dawei, et al.
Published: (2025)
MASCOT: Towards Multi-Agent Socio-Collaborative Companion Systems
by: Wang, Yiyang, et al.
Published: (2026)
by: Wang, Yiyang, et al.
Published: (2026)
Doc-V*:Coarse-to-Fine Interactive Visual Reasoning for Multi-Page Document VQA
by: Zheng, Yuanlei, et al.
Published: (2026)
by: Zheng, Yuanlei, et al.
Published: (2026)
TASER: Table Agents for Schema-guided Extraction and Recommendation
by: Cho, Nicole, et al.
Published: (2025)
by: Cho, Nicole, et al.
Published: (2025)
What Makes a Good Query? Measuring the Impact of Human-Confusing Linguistic Features on LLM Performance
by: Watson, William, et al.
Published: (2026)
by: Watson, William, et al.
Published: (2026)
ScreenLLM: Stateful Screen Schema for Efficient Action Understanding and Prediction
by: Jin, Yiqiao, et al.
Published: (2025)
by: Jin, Yiqiao, et al.
Published: (2025)
SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression
by: Jin, Yiqiao, et al.
Published: (2025)
by: Jin, Yiqiao, et al.
Published: (2025)
Sysformer: Safeguarding Frozen Large Language Models with Adaptive System Prompts
by: Sharma, Kartik, et al.
Published: (2025)
by: Sharma, Kartik, et al.
Published: (2025)
UniGuard: Towards Universal Safety Guardrails for Jailbreak Attacks on Multimodal Large Language Models
by: Oh, Sejoon, et al.
Published: (2024)
by: Oh, Sejoon, et al.
Published: (2024)
Turning Logic Against Itself : Probing Model Defenses Through Contrastive Questions
by: Sachdeva, Rachneet, et al.
Published: (2025)
by: Sachdeva, Rachneet, et al.
Published: (2025)
CATfOOD: Counterfactual Augmented Training for Improving Out-of-Domain Performance and Calibration
by: Sachdeva, Rachneet, et al.
Published: (2023)
by: Sachdeva, Rachneet, et al.
Published: (2023)
Learning and Calibrating Heterogeneous Bounded Rational Market Behaviour with Multi-Agent Reinforcement Learning
by: Evans, Benjamin Patrick, et al.
Published: (2024)
by: Evans, Benjamin Patrick, et al.
Published: (2024)
Multi-Agent Interactive Question Generation Framework for Long Document Understanding
by: Wang, Kesen, et al.
Published: (2025)
by: Wang, Kesen, et al.
Published: (2025)
Continual Learning of Domain Knowledge from Human Feedback in Text-to-SQL
by: Cook, Thomas, et al.
Published: (2025)
by: Cook, Thomas, et al.
Published: (2025)
UniSD: Towards a Unified Self-Distillation Framework for Large Language Models
by: Jin, Yiqiao, et al.
Published: (2026)
by: Jin, Yiqiao, et al.
Published: (2026)
AgentSOC: A Multi-Layer Agentic AI Framework for Security Operations Automation
by: Roy, Joyjit, et al.
Published: (2026)
by: Roy, Joyjit, et al.
Published: (2026)
CompeteAI: Understanding the Competition Dynamics in Large Language Model-based Agents
by: Zhao, Qinlin, et al.
Published: (2023)
by: Zhao, Qinlin, et al.
Published: (2023)
Localizing and Mitigating Errors in Long-form Question Answering
by: Sachdeva, Rachneet, et al.
Published: (2024)
by: Sachdeva, Rachneet, et al.
Published: (2024)
AgentReview: Exploring Peer Review Dynamics with LLM Agents
by: Jin, Yiqiao, et al.
Published: (2024)
by: Jin, Yiqiao, et al.
Published: (2024)
QueryBandits for Hallucination Mitigation: Exploiting Semantic Features for No-Regret Rewriting
by: Cho, Nicole, et al.
Published: (2025)
by: Cho, Nicole, et al.
Published: (2025)
No One Size Fits All: QueryBandits for Hallucination Mitigation
by: Cho, Nicole, et al.
Published: (2026)
by: Cho, Nicole, et al.
Published: (2026)
BEAVER: A Training-Free Hierarchical Prompt Compression Method via Structure-Aware Page Selection
by: Hu, Zhengpei, et al.
Published: (2026)
by: Hu, Zhengpei, et al.
Published: (2026)
From Pixels to Predictions: Spectrogram and Vision Transformer for Better Time Series Forecasting
by: Zeng, Zhen, et al.
Published: (2024)
by: Zeng, Zhen, et al.
Published: (2024)
Memory-Augmented Agent Training for Business Document Understanding
by: Liu, Jiale, et al.
Published: (2024)
by: Liu, Jiale, et al.
Published: (2024)
GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
by: Xu, Yuancheng, et al.
Published: (2024)
by: Xu, Yuancheng, et al.
Published: (2024)
HiMATE: A Hierarchical Multi-Agent Framework for Machine Translation Evaluation
by: Zhang, Shijie, et al.
Published: (2025)
by: Zhang, Shijie, et al.
Published: (2025)
Multi2: Multi-Agent Test-Time Scalable Framework for Multi-Document Processing
by: Cao, Juntai, et al.
Published: (2025)
by: Cao, Juntai, et al.
Published: (2025)
TopoChunker: Topology-Aware Agentic Document Chunking Framework
by: Liu, Xiaoyu
Published: (2026)
by: Liu, Xiaoyu
Published: (2026)
Remember Your Trace: Memory-Guided Long-Horizon Agentic Framework for Consistent and Hierarchical Repository-Level Code Documentation
by: Bae, Suyoung, et al.
Published: (2026)
by: Bae, Suyoung, et al.
Published: (2026)
HiRAS: A Hierarchical Multi-Agent Framework for Paper-to-Code Generation and Execution
by: Hong, Hanhua, et al.
Published: (2026)
by: Hong, Hanhua, et al.
Published: (2026)
Similar Items
-
ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering
by: Kaur, Rachneet, et al.
Published: (2025) -
AdaptAgent: Adapting Multimodal Web Agents with Few-Shot Learning from Human Demonstrations
by: Verma, Gaurav, et al.
Published: (2024) -
Evaluating Large Language Models on Time Series Feature Understanding: A Comprehensive Taxonomy and Benchmark
by: Fons, Elizabeth, et al.
Published: (2024) -
LETS-C: Leveraging Text Embedding for Time Series Classification
by: Kaur, Rachneet, et al.
Published: (2024) -
Efficient Knowledge Probing of Large Language Models by Adapting Pre-trained Embeddings
by: Sharma, Kartik, et al.
Published: (2025)