FormFactory: An Interactive Benchmarking Suite for Multimodal Form-Filling Agents
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Bobo, Wang, Yuheng, Fei, Hao, Li, Juncheng, Ji, Wei, Lee, Mong-Li, Hsu, Wynne |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis
by: Luo, Meng, et al.
Published: (2024)
by: Luo, Meng, et al.
Published: (2024)
Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment
by: Li, Bobo, et al.
Published: (2026)
by: Li, Bobo, et al.
Published: (2026)
Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
by: Yan, Zehong, et al.
Published: (2025)
by: Yan, Zehong, et al.
Published: (2025)
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
by: Qi, Peng, et al.
Published: (2024)
by: Qi, Peng, et al.
Published: (2024)
ChronoFact: Timeline-based Temporal Fact Verification
by: Barik, Anab Maulana, et al.
Published: (2024)
by: Barik, Anab Maulana, et al.
Published: (2024)
Faithful Logical Reasoning via Symbolic Chain-of-Thought
by: Xu, Jundong, et al.
Published: (2024)
by: Xu, Jundong, et al.
Published: (2024)
From Personas to Talks: Revisiting the Impact of Personas on LLM-Synthesized Emotional Support Conversations
by: Wu, Shenghan, et al.
Published: (2025)
by: Wu, Shenghan, et al.
Published: (2025)
Orthogonal Spatial-temporal Distributional Transfer for 4D Generation
by: Liu, Wei, et al.
Published: (2026)
by: Liu, Wei, et al.
Published: (2026)
When Disagreements Elicit Robustness: Investigating Self-Repair Capabilities under LLM Multi-Agent Disagreements
by: Ju, Tianjie, et al.
Published: (2025)
by: Ju, Tianjie, et al.
Published: (2025)
CMNER: A Chinese Multimodal NER Dataset based on Social Media
by: Ji, Yuanze, et al.
Published: (2024)
by: Ji, Yuanze, et al.
Published: (2024)
On the Adaptive Psychological Persuasion of Large Language Models
by: Ju, Tianjie, et al.
Published: (2025)
by: Ju, Tianjie, et al.
Published: (2025)
LogicReward: Incentivizing LLM Reasoning via Step-Wise Logical Supervision
by: Xu, Jundong, et al.
Published: (2025)
by: Xu, Jundong, et al.
Published: (2025)
Aristotle: Mastering Logical Reasoning with A Logic-Complete Decompose-Search-Resolve Framework
by: Xu, Jundong, et al.
Published: (2024)
by: Xu, Jundong, et al.
Published: (2024)
Probing then Editing Response Personality of Large Language Models
by: Ju, Tianjie, et al.
Published: (2025)
by: Ju, Tianjie, et al.
Published: (2025)
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction
by: Liu, Jiang, et al.
Published: (2024)
by: Liu, Jiang, et al.
Published: (2024)
TRUST-VL: An Explainable News Assistant for General Multimodal Misinformation Detection
by: Yan, Zehong, et al.
Published: (2025)
by: Yan, Zehong, et al.
Published: (2025)
Multi-Part Object Representations via Graph Structures and Co-Part Discovery
by: Foo, Alex, et al.
Published: (2025)
by: Foo, Alex, et al.
Published: (2025)
QIME: Constructing Interpretable Medical Text Embeddings via Ontology-Grounded Questions
by: Tang, Yixuan, et al.
Published: (2026)
by: Tang, Yixuan, et al.
Published: (2026)
Althea: Human-AI Collaboration for Fact-Checking and Critical Reasoning
by: Churina, Svetlana, et al.
Published: (2025)
by: Churina, Svetlana, et al.
Published: (2025)
Watch Out Your Album! On the Inadvertent Privacy Memorization in Multi-Modal Large Language Models
by: Ju, Tianjie, et al.
Published: (2025)
by: Ju, Tianjie, et al.
Published: (2025)
Revisiting Structured Sentiment Analysis as Latent Dependency Graph Parsing
by: Zhou, Chengjie, et al.
Published: (2024)
by: Zhou, Chengjie, et al.
Published: (2024)
Unveiling the Cognitive Compass: Theory-of-Mind-Guided Multimodal Emotion Reasoning
by: Luo, Meng, et al.
Published: (2026)
by: Luo, Meng, et al.
Published: (2026)
LEAF-Mamba: Local Emphatic and Adaptive Fusion State Space Model for RGB-D Salient Object Detection
by: Wu, Lanhu, et al.
Published: (2025)
by: Wu, Lanhu, et al.
Published: (2025)
XFormParser: A Simple and Effective Multimodal Multilingual Semi-structured Form Parser
by: Cheng, Xianfu, et al.
Published: (2024)
by: Cheng, Xianfu, et al.
Published: (2024)
NUS-Emo at SemEval-2024 Task 3: Instruction-Tuning LLM for Multimodal Emotion-Cause Analysis in Conversations
by: Luo, Meng, et al.
Published: (2024)
by: Luo, Meng, et al.
Published: (2024)
Video-of-Thought: Step-by-Step Video Reasoning from Perception to Cognition
by: Fei, Hao, et al.
Published: (2024)
by: Fei, Hao, et al.
Published: (2024)
LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
by: Wu, Yuhao, et al.
Published: (2024)
by: Wu, Yuhao, et al.
Published: (2024)
Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
by: Chen, Junjie, et al.
Published: (2026)
by: Chen, Junjie, et al.
Published: (2026)
KnowMT-Bench: Benchmarking Knowledge-Intensive Long-Form Question Answering in Multi-Turn Dialogues
by: Chen, Junhao, et al.
Published: (2025)
by: Chen, Junhao, et al.
Published: (2025)
Evaluating the Robustness of Multimodal Agents Against Active Environmental Injection Attacks
by: Chen, Yurun, et al.
Published: (2025)
by: Chen, Yurun, et al.
Published: (2025)
SuperWriter: Reflection-Driven Long-Form Generation with Large Language Models
by: Wu, Yuhao, et al.
Published: (2025)
by: Wu, Yuhao, et al.
Published: (2025)
UniM: A Unified Any-to-Any Interleaved Multimodal Benchmark
by: Li, Yanlin, et al.
Published: (2026)
by: Li, Yanlin, et al.
Published: (2026)
SEACrowd: A Multilingual Multimodal Data Hub and Benchmark Suite for Southeast Asian Languages
by: Lovenia, Holy, et al.
Published: (2024)
by: Lovenia, Holy, et al.
Published: (2024)
LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
by: Yuan, Peiwen, et al.
Published: (2025)
by: Yuan, Peiwen, et al.
Published: (2025)
Benchmarking Multimodal Retrieval Augmented Generation with Dynamic VQA Dataset and Self-adaptive Planning Agent
by: Li, Yangning, et al.
Published: (2024)
by: Li, Yangning, et al.
Published: (2024)
Precise Information Control in Long-Form Text Generation
by: He, Jacqueline, et al.
Published: (2025)
by: He, Jacqueline, et al.
Published: (2025)
UNCLE: Benchmarking Uncertainty Expressions in Long-Form Generation
by: Yang, Ruihan, et al.
Published: (2025)
by: Yang, Ruihan, et al.
Published: (2025)
AnesSuite: A Comprehensive Benchmark and Dataset Suite for Anesthesiology Reasoning in LLMs
by: Feng, Xiang, et al.
Published: (2025)
by: Feng, Xiang, et al.
Published: (2025)
Multi-Granular Multimodal Clue Fusion for Meme Understanding
by: Zheng, Li, et al.
Published: (2025)
by: Zheng, Li, et al.
Published: (2025)
Dynamic Emotion and Personality Profiling for Multimodal Deception Detection
by: Zheng, Li, et al.
Published: (2026)
by: Zheng, Li, et al.
Published: (2026)
Similar Items
-
PanoSent: A Panoptic Sextuple Extraction Benchmark for Multimodal Conversational Aspect-based Sentiment Analysis
by: Luo, Meng, et al.
Published: (2024) -
Taming Actor-Observer Asymmetry in Agents via Dialectical Alignment
by: Li, Bobo, et al.
Published: (2026) -
Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
by: Yan, Zehong, et al.
Published: (2025) -
SNIFFER: Multimodal Large Language Model for Explainable Out-of-Context Misinformation Detection
by: Qi, Peng, et al.
Published: (2024) -
ChronoFact: Timeline-based Temporal Fact Verification
by: Barik, Anab Maulana, et al.
Published: (2024)