IsoBench: Benchmarking Multimodal Foundation Models on Isomorphic Representations
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Fu, Deqing, Guo, Ruohao, Khalighinejad, Ghazal, Liu, Ollie, Dhingra, Bhuwan, Yogatama, Dani, Jia, Robin, Neiswanger, Willie |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
DeLLMa: Decision Making Under Uncertainty with Large Language Models
von: Liu, Ollie, et al.
Veröffentlicht: (2024)
von: Liu, Ollie, et al.
Veröffentlicht: (2024)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025)
Document-as-Image Representations Fall Short for Scientific Retrieval
von: Khalighinejad, Ghazal, et al.
Veröffentlicht: (2026)
von: Khalighinejad, Ghazal, et al.
Veröffentlicht: (2026)
MatViX: Multimodal Information Extraction from Visually Rich Articles
von: Khalighinejad, Ghazal, et al.
Veröffentlicht: (2024)
von: Khalighinejad, Ghazal, et al.
Veröffentlicht: (2024)
Extracting Polymer Nanocomposite Samples from Full-Length Documents
von: Khalighinejad, Ghazal, et al.
Veröffentlicht: (2024)
von: Khalighinejad, Ghazal, et al.
Veröffentlicht: (2024)
LLM Unlearning Without an Expert Curated Dataset
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
von: Zhu, Xiaoyuan, et al.
Veröffentlicht: (2025)
Resa: Transparent Reasoning Models via SAEs
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)
Euclid: Supercharging Multimodal LLMs with Synthetic High-Fidelity Visual Descriptions
von: Zhang, Jiarui, et al.
Veröffentlicht: (2024)
von: Zhang, Jiarui, et al.
Veröffentlicht: (2024)
METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring
von: Liu, Ollie, et al.
Veröffentlicht: (2025)
von: Liu, Ollie, et al.
Veröffentlicht: (2025)
Pelican Soup Framework: A Theoretical Framework for Language Model Capabilities
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2024)
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2024)
On Retrieval Augmentation and the Limitations of Language Model Training
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2023)
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2023)
GenEOL: Harnessing the Generative Power of LLMs for Training-Free Sentence Embeddings
von: Thirukovalluru, Raghuveer, et al.
Veröffentlicht: (2024)
von: Thirukovalluru, Raghuveer, et al.
Veröffentlicht: (2024)
The Rotary Position Embedding May Cause Dimension Inefficiency in Attention Heads for Long-Distance Retrieval
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2025)
von: Chiang, Ting-Rui, et al.
Veröffentlicht: (2025)
Precise Debugging Benchmark: Is Your Model Debugging or Regenerating?
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026)
von: Zhu, Wang Bill, et al.
Veröffentlicht: (2026)
Tina: Tiny Reasoning Models via LoRA
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)
von: Wang, Shangshang, et al.
Veröffentlicht: (2025)
Raccoon: Prompt Extraction Benchmark of LLM-Integrated Applications
von: Wang, Junlin, et al.
Veröffentlicht: (2024)
von: Wang, Junlin, et al.
Veröffentlicht: (2024)
The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities
von: Wu, Zhaofeng, et al.
Veröffentlicht: (2024)
von: Wu, Zhaofeng, et al.
Veröffentlicht: (2024)
Atomic Self-Consistency for Better Long Form Generations
von: Thirukovalluru, Raghuveer, et al.
Veröffentlicht: (2024)
von: Thirukovalluru, Raghuveer, et al.
Veröffentlicht: (2024)
ChatShop: Interactive Information Seeking with Language Agents
von: Chen, Sanxing, et al.
Veröffentlicht: (2024)
von: Chen, Sanxing, et al.
Veröffentlicht: (2024)
Real-time Factuality Assessment from Adversarial Feedback
von: Chen, Sanxing, et al.
Veröffentlicht: (2024)
von: Chen, Sanxing, et al.
Veröffentlicht: (2024)
RVPO: Risk-Sensitive Alignment via Variance Regularization
von: Montero, Ivan, et al.
Veröffentlicht: (2026)
von: Montero, Ivan, et al.
Veröffentlicht: (2026)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
von: Li, Ang, et al.
Veröffentlicht: (2025)
von: Li, Ang, et al.
Veröffentlicht: (2025)
To Trust or Not to Trust? Enhancing Large Language Models' Situated Faithfulness to External Contexts
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
Automated Benchmark Auditing for AI Agents and Large Language Models
von: Wang, Junlin, et al.
Veröffentlicht: (2026)
von: Wang, Junlin, et al.
Veröffentlicht: (2026)
SCENE: Self-Labeled Counterfactuals for Extrapolating to Negative Examples
von: Fu, Deqing, et al.
Veröffentlicht: (2023)
von: Fu, Deqing, et al.
Veröffentlicht: (2023)
Convergent Evolution: How Different Language Models Learn Similar Number Representations
von: Fu, Deqing, et al.
Veröffentlicht: (2026)
von: Fu, Deqing, et al.
Veröffentlicht: (2026)
Causal Interventions on Causal Paths: Mapping GPT-2's Reasoning From Syntax to Semantics
von: Lee, Isabelle, et al.
Veröffentlicht: (2024)
von: Lee, Isabelle, et al.
Veröffentlicht: (2024)
Hierarchical Multi-Label Classification of Online Vaccine Concerns
von: Zhu, Chloe Qinyu, et al.
Veröffentlicht: (2024)
von: Zhu, Chloe Qinyu, et al.
Veröffentlicht: (2024)
Rethinking RL for LLM Reasoning: It's Sparse Policy Selection, Not Capability Learning
von: Akgül, Ömer Faruk, et al.
Veröffentlicht: (2026)
von: Akgül, Ömer Faruk, et al.
Veröffentlicht: (2026)
Adversarial Math Word Problem Generation
von: Xie, Roy, et al.
Veröffentlicht: (2024)
von: Xie, Roy, et al.
Veröffentlicht: (2024)
Coding Agents are Effective Long-Context Processors
von: Cao, Weili, et al.
Veröffentlicht: (2026)
von: Cao, Weili, et al.
Veröffentlicht: (2026)
MasalBench: A Benchmark for Contextual and Cross-Cultural Understanding of Persian Proverbs in LLMs
von: Kalhor, Ghazal, et al.
Veröffentlicht: (2026)
von: Kalhor, Ghazal, et al.
Veröffentlicht: (2026)
Pre-trained Large Language Models Use Fourier Features to Compute Addition
von: Zhou, Tianyi, et al.
Veröffentlicht: (2024)
von: Zhou, Tianyi, et al.
Veröffentlicht: (2024)
A Platform for Investigating Public Health Content with Efficient Concern Classification
von: Li, Christopher, et al.
Veröffentlicht: (2025)
von: Li, Christopher, et al.
Veröffentlicht: (2025)
From Calibration to Collaboration: LLM Uncertainty Quantification Should Be More Human-Centered
von: Devic, Siddartha, et al.
Veröffentlicht: (2025)
von: Devic, Siddartha, et al.
Veröffentlicht: (2025)
Calibrating Long-form Generations from Large Language Models
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
von: Huang, Yukun, et al.
Veröffentlicht: (2024)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
von: Huang, Yukun, et al.
Veröffentlicht: (2025)
von: Huang, Yukun, et al.
Veröffentlicht: (2025)
InData: Towards Secure Multi-Step, Tool-Based Data Analysis
von: K, Karthikeyan, et al.
Veröffentlicht: (2025)
von: K, Karthikeyan, et al.
Veröffentlicht: (2025)
TokenSmith: Streamlining Data Editing, Search, and Inspection for Large-Scale Language Model Training and Interpretability
von: Khan, Mohammad Aflah, et al.
Veröffentlicht: (2025)
von: Khan, Mohammad Aflah, et al.
Veröffentlicht: (2025)
Atomic Consistency Preference Optimization for Long-Form Question Answering
von: Chen, Jingfeng, et al.
Veröffentlicht: (2025)
von: Chen, Jingfeng, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
DeLLMa: Decision Making Under Uncertainty with Large Language Models
von: Liu, Ollie, et al.
Veröffentlicht: (2024) -
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
von: Gan, Woody Haosheng, et al.
Veröffentlicht: (2025) -
Document-as-Image Representations Fall Short for Scientific Retrieval
von: Khalighinejad, Ghazal, et al.
Veröffentlicht: (2026) -
MatViX: Multimodal Information Extraction from Visually Rich Articles
von: Khalighinejad, Ghazal, et al.
Veröffentlicht: (2024) -
Extracting Polymer Nanocomposite Samples from Full-Length Documents
von: Khalighinejad, Ghazal, et al.
Veröffentlicht: (2024)