M-Longdoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning Framework
Fuente:
arXiv
Salvato in:
| Autori principali: | Chia, Yew Ken, Cheng, Liying, Chan, Hou Pong, Liu, Chaoqun, Song, Maojia, Aljunied, Sharifah Mahani, Poria, Soujanya, Bing, Lidong |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Domain-Expanded ASTE: Rethinking Generalization in Aspect Sentiment Triplet Extraction
di: Chia, Yew Ken, et al.
Pubblicazione: (2023)
di: Chia, Yew Ken, et al.
Pubblicazione: (2023)
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
di: Chia, Yew Ken, et al.
Pubblicazione: (2024)
di: Chia, Yew Ken, et al.
Pubblicazione: (2024)
Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse Prompts
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2023)
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2023)
PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
di: Chia, Yew Ken, et al.
Pubblicazione: (2024)
di: Chia, Yew Ken, et al.
Pubblicazione: (2024)
SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages
di: Zhang, Wenxuan, et al.
Pubblicazione: (2024)
di: Zhang, Wenxuan, et al.
Pubblicazione: (2024)
SeaExam and SeaBench: Benchmarking LLMs with Local Multilingual Questions in Southeast Asia
di: Liu, Chaoqun, et al.
Pubblicazione: (2025)
di: Liu, Chaoqun, et al.
Pubblicazione: (2025)
Analyzing LLMs' Knowledge Boundary Cognition Across Languages Through the Lens of Internal Representations
di: Xiao, Chenghao, et al.
Pubblicazione: (2025)
di: Xiao, Chenghao, et al.
Pubblicazione: (2025)
Reasoning Paths Optimization: Learning to Reason and Explore From Diverse Paths
di: Chia, Yew Ken, et al.
Pubblicazione: (2024)
di: Chia, Yew Ken, et al.
Pubblicazione: (2024)
Chain-of-Knowledge: Grounding Large Language Models via Dynamic Knowledge Adapting over Heterogeneous Sources
di: Li, Xingxuan, et al.
Pubblicazione: (2023)
di: Li, Xingxuan, et al.
Pubblicazione: (2023)
SeaLLMs-Audio: Large Audio-Language Models for Southeast Asia
di: Liu, Chaoqun, et al.
Pubblicazione: (2025)
di: Liu, Chaoqun, et al.
Pubblicazione: (2025)
Are Language Models Puzzle Prodigies? Algorithmic Puzzles Unveil Serious Challenges in Multimodal Reasoning
di: Ghosal, Deepanway, et al.
Pubblicazione: (2024)
di: Ghosal, Deepanway, et al.
Pubblicazione: (2024)
The Jumping Reasoning Curve? Tracking the Evolution of Reasoning Performance in GPT-[n] and o-[n] Models on Multimodal Puzzles
di: Toh, Vernon Y. H., et al.
Pubblicazione: (2025)
di: Toh, Vernon Y. H., et al.
Pubblicazione: (2025)
PromptDistill: Query-based Selective Token Retention in Intermediate Layers for Efficient Large Language Model Inference
di: Jin, Weisheng, et al.
Pubblicazione: (2025)
di: Jin, Weisheng, et al.
Pubblicazione: (2025)
Babel: Open Multilingual Large Language Models Serving Over 90% of Global Speakers
di: Zhao, Yiran, et al.
Pubblicazione: (2025)
di: Zhao, Yiran, et al.
Pubblicazione: (2025)
Scaling Language-Centric Omnimodal Representation Learning
di: Xiao, Chenghao, et al.
Pubblicazione: (2025)
di: Xiao, Chenghao, et al.
Pubblicazione: (2025)
Towards Robust Instruction Tuning on Multimodal Large Language Models
di: Han, Wei, et al.
Pubblicazione: (2024)
di: Han, Wei, et al.
Pubblicazione: (2024)
SeaLLMs -- Large Language Models for Southeast Asia
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2023)
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2023)
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
di: LASA Team, et al.
Pubblicazione: (2025)
di: LASA Team, et al.
Pubblicazione: (2025)
Video2Music: Suitable Music Generation from Videos using an Affective Multimodal Transformer model
di: Kang, Jaeyong, et al.
Pubblicazione: (2023)
di: Kang, Jaeyong, et al.
Pubblicazione: (2023)
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
di: Song, Maojia, et al.
Pubblicazione: (2024)
di: Song, Maojia, et al.
Pubblicazione: (2024)
Inference Time Alignment with Reward-Guided Tree Search
di: Hung, Chia-Yu, et al.
Pubblicazione: (2024)
di: Hung, Chia-Yu, et al.
Pubblicazione: (2024)
Toward Robust Multimodal Learning using Multimodal Foundational Models
di: Zhao, Xianbing, et al.
Pubblicazione: (2024)
di: Zhao, Xianbing, et al.
Pubblicazione: (2024)
Consistency Guided Knowledge Retrieval and Denoising in LLMs for Zero-shot Document-level Relation Triplet Extraction
di: Sun, Qi, et al.
Pubblicazione: (2024)
di: Sun, Qi, et al.
Pubblicazione: (2024)
Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions
di: Zhao, Ruochen, et al.
Pubblicazione: (2024)
di: Zhao, Ruochen, et al.
Pubblicazione: (2024)
Exact Flow Linear Attention: Exact Solution from Continuous-Time Dynamics
di: Lei, Jingdi, et al.
Pubblicazione: (2025)
di: Lei, Jingdi, et al.
Pubblicazione: (2025)
PREMISE: Matching-based Prediction for Accurate Review Recommendation
di: Han, Wei, et al.
Pubblicazione: (2025)
di: Han, Wei, et al.
Pubblicazione: (2025)
FINEREASON: Evaluating and Improving LLMs' Deliberate Reasoning through Reflective Puzzle Solving
di: Chen, Guizhen, et al.
Pubblicazione: (2025)
di: Chen, Guizhen, et al.
Pubblicazione: (2025)
LLMs Can't Handle Peer Pressure: Crumbling under Multi-Agent Social Interactions
di: Song, Maojia, et al.
Pubblicazione: (2025)
di: Song, Maojia, et al.
Pubblicazione: (2025)
From Perception to Action: An Interactive Benchmark for Vision Reasoning
di: Wu, Yuhao, et al.
Pubblicazione: (2026)
di: Wu, Yuhao, et al.
Pubblicazione: (2026)
Demystifying deep search: a holistic evaluation with hint-free multi-hop questions and factorised metrics
di: Song, Maojia, et al.
Pubblicazione: (2025)
di: Song, Maojia, et al.
Pubblicazione: (2025)
Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
di: Bhardwaj, Rishabh, et al.
Pubblicazione: (2024)
di: Bhardwaj, Rishabh, et al.
Pubblicazione: (2024)
DELLA-Merging: Reducing Interference in Model Merging through Magnitude-Based Sampling
di: Deep, Pala Tej, et al.
Pubblicazione: (2024)
di: Deep, Pala Tej, et al.
Pubblicazione: (2024)
MMDocIR: Benchmarking Multimodal Retrieval for Long Documents
di: Dong, Kuicai, et al.
Pubblicazione: (2025)
di: Dong, Kuicai, et al.
Pubblicazione: (2025)
Understanding the Capabilities and Limitations of Large Language Models for Cultural Commonsense
di: Shen, Siqi, et al.
Pubblicazione: (2024)
di: Shen, Siqi, et al.
Pubblicazione: (2024)
Sufi Warriorism in Muslim Southeast Asia
di: Khairudin Aljunied
Pubblicazione: (2024)
di: Khairudin Aljunied
Pubblicazione: (2024)
Stacked from One: Multi-Scale Self-Injection for Context Window Extension
di: Han, Wei, et al.
Pubblicazione: (2026)
di: Han, Wei, et al.
Pubblicazione: (2026)
Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming
di: Han, Vernon Toh Yan, et al.
Pubblicazione: (2024)
di: Han, Vernon Toh Yan, et al.
Pubblicazione: (2024)
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control
di: Zhang, Shaozuo, et al.
Pubblicazione: (2025)
di: Zhang, Shaozuo, et al.
Pubblicazione: (2025)
Sowing the Wind, Reaping the Whirlwind: The Impact of Editing Language Models
di: Hazra, Rima, et al.
Pubblicazione: (2024)
di: Hazra, Rima, et al.
Pubblicazione: (2024)
Safety Arithmetic: A Framework for Test-time Safety Alignment of Language Models by Steering Parameters and Activations
di: Hazra, Rima, et al.
Pubblicazione: (2024)
di: Hazra, Rima, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Domain-Expanded ASTE: Rethinking Generalization in Aspect Sentiment Triplet Extraction
di: Chia, Yew Ken, et al.
Pubblicazione: (2023) -
Can-Do! A Dataset and Neuro-Symbolic Grounded Framework for Embodied Planning with Large Multimodal Models
di: Chia, Yew Ken, et al.
Pubblicazione: (2024) -
Democratizing LLMs for Low-Resource Languages by Leveraging their English Dominant Abilities with Linguistically-Diverse Prompts
di: Nguyen, Xuan-Phi, et al.
Pubblicazione: (2023) -
PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns
di: Chia, Yew Ken, et al.
Pubblicazione: (2024) -
SeaLLMs 3: Open Foundation and Chat Multilingual Large Language Models for Southeast Asian Languages
di: Zhang, Wenxuan, et al.
Pubblicazione: (2024)