Eureka: Evaluating and Understanding Large Foundation Models
Fuente:
arXiv
Saved in:
| Main Authors: | Balachandran, Vidhisha, Chen, Jingya, Joshi, Neel, Nushi, Besmira, Palangi, Hamid, Salinas, Eduardo, Vineet, Vibhav, Woffinden-Luey, James, Yousefi, Safoora |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead
by: Balachandran, Vidhisha, et al.
Published: (2025)
by: Balachandran, Vidhisha, et al.
Published: (2025)
Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models
by: Moayeri, Mazda, et al.
Published: (2024)
by: Moayeri, Mazda, et al.
Published: (2024)
MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation
by: Joshi, Siddharth, et al.
Published: (2025)
by: Joshi, Siddharth, et al.
Published: (2025)
Improving Instruction-Following in Language Models through Activation Steering
by: Stolfo, Alessandro, et al.
Published: (2024)
by: Stolfo, Alessandro, et al.
Published: (2024)
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
by: Butt, Natasha, et al.
Published: (2024)
by: Butt, Natasha, et al.
Published: (2024)
Tracing the Traces: Latent Temporal Signals for Efficient and Accurate Reasoning
by: Vilas, Martina G., et al.
Published: (2025)
by: Vilas, Martina G., et al.
Published: (2025)
Just Do It!? Computer-Use Agents Exhibit Blind Goal-Directedness
by: Shayegani, Erfan, et al.
Published: (2025)
by: Shayegani, Erfan, et al.
Published: (2025)
Diversity of Thought Improves Reasoning Abilities of LLMs
by: Naik, Ranjita, et al.
Published: (2023)
by: Naik, Ranjita, et al.
Published: (2023)
Detecting Data Contamination in LLMs via In-Context Learning
by: Zawalski, Michał, et al.
Published: (2025)
by: Zawalski, Michał, et al.
Published: (2025)
Phi-4-reasoning Technical Report
by: Abdin, Marah, et al.
Published: (2025)
by: Abdin, Marah, et al.
Published: (2025)
Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models
by: Yuksekgonul, Mert, et al.
Published: (2023)
by: Yuksekgonul, Mert, et al.
Published: (2023)
What MLLMs Learn about When they Learn about Multimodal Reasoning
by: Chung, Jiwan, et al.
Published: (2025)
by: Chung, Jiwan, et al.
Published: (2025)
Attention Speaks Volumes: Localizing and Mitigating Bias in Language Models
by: Adiga, Rishabh, et al.
Published: (2024)
by: Adiga, Rishabh, et al.
Published: (2024)
Understanding Information Storage and Transfer in Multi-modal Large Language Models
by: Basu, Samyadeep, et al.
Published: (2024)
by: Basu, Samyadeep, et al.
Published: (2024)
MMMT-IF: A Challenging Multimodal Multi-Turn Instruction Following Benchmark
by: Epstein, Elliot L., et al.
Published: (2024)
by: Epstein, Elliot L., et al.
Published: (2024)
Exploring Group and Symmetry Principles in Large Language Models
by: Imani, Shima, et al.
Published: (2024)
by: Imani, Shima, et al.
Published: (2024)
Elephants Never Forget: Memorization and Learning of Tabular Data in Large Language Models
by: Bordt, Sebastian, et al.
Published: (2024)
by: Bordt, Sebastian, et al.
Published: (2024)
KodeXv0.1: A Family of State-of-the-Art Financial Large Language Models
by: Rajani, Neel, et al.
Published: (2024)
by: Rajani, Neel, et al.
Published: (2024)
Understanding Depth and Height Perception in Large Visual-Language Models
by: Azad, Shehreen, et al.
Published: (2024)
by: Azad, Shehreen, et al.
Published: (2024)
Physics Knowledge in Frontier Models: A Diagnostic Study of Failure Modes
by: Bagdonaviciute, Ieva, et al.
Published: (2025)
by: Bagdonaviciute, Ieva, et al.
Published: (2025)
Value Lens: Using Large Language Models to Understand Human Values
by: Fernández, Eduardo de la Cruz, et al.
Published: (2025)
by: Fernández, Eduardo de la Cruz, et al.
Published: (2025)
Evaluating Foundation Models' 3D Understanding Through Multi-View Correspondence Analysis
by: Lilova, Valentina, et al.
Published: (2025)
by: Lilova, Valentina, et al.
Published: (2025)
HierarQ: Task-Aware Hierarchical Q-Former for Enhanced Video Understanding
by: Azad, Shehreen, et al.
Published: (2025)
by: Azad, Shehreen, et al.
Published: (2025)
Foundation Models for Geospatial Reasoning: Assessing Capabilities of Large Language Models in Understanding Geometries and Topological Spatial Relations
by: Ji, Yuhan, et al.
Published: (2025)
by: Ji, Yuhan, et al.
Published: (2025)
Decoding In-Context Learning: Neuroscience-inspired Analysis of Representations in Large Language Models
by: Yousefi, Safoora, et al.
Published: (2023)
by: Yousefi, Safoora, et al.
Published: (2023)
Is A Picture Worth A Thousand Words? Delving Into Spatial Reasoning for Vision Language Models
by: Wang, Jiayu, et al.
Published: (2024)
by: Wang, Jiayu, et al.
Published: (2024)
KGQuiz: Evaluating the Generalization of Encoded Knowledge in Large Language Models
by: Bai, Yuyang, et al.
Published: (2023)
by: Bai, Yuyang, et al.
Published: (2023)
Exposing the Achilles' Heel: Evaluating LLMs Ability to Handle Mistakes in Mathematical Reasoning
by: Singh, Joykirat, et al.
Published: (2024)
by: Singh, Joykirat, et al.
Published: (2024)
Emotion-Attended Stateful Memory (EASM):The Architecture for Hyper-Personalization at Scale
by: Kotecha, Vineet, et al.
Published: (2026)
by: Kotecha, Vineet, et al.
Published: (2026)
Evaluating Large Language Models for Causal Modeling
by: Razouk, Houssam, et al.
Published: (2024)
by: Razouk, Houssam, et al.
Published: (2024)
Do LLMs Use Cultural Knowledge Without Being Told? A Multilingual Evaluation of Implicit Pragmatic Adaptation
by: Nasim, Mehwish, et al.
Published: (2026)
by: Nasim, Mehwish, et al.
Published: (2026)
CORDIAL: Can Multimodal Large Language Models Effectively Understand Coherence Relationships?
by: Ramakrishnan, Aashish Anantha, et al.
Published: (2025)
by: Ramakrishnan, Aashish Anantha, et al.
Published: (2025)
Understanding Gen Alpha Digital Language: Evaluation of LLM Safety Systems for Content Moderation
by: Mehta, Manisha, et al.
Published: (2025)
by: Mehta, Manisha, et al.
Published: (2025)
Towards Human Understanding of Paraphrase Types in Large Language Models
by: Meier, Dominik, et al.
Published: (2024)
by: Meier, Dominik, et al.
Published: (2024)
A Large-Scale Analysis on Contextual Self-Supervised Video Representation Learning
by: Kumar, Akash, et al.
Published: (2025)
by: Kumar, Akash, et al.
Published: (2025)
TEncDM: Understanding the Properties of the Diffusion Model in the Space of Language Model Encodings
by: Shabalin, Alexander, et al.
Published: (2024)
by: Shabalin, Alexander, et al.
Published: (2024)
Large Language Models Can Better Understand Knowledge Graphs Than We Thought
by: Dai, Xinbang, et al.
Published: (2024)
by: Dai, Xinbang, et al.
Published: (2024)
Navigating Hallucinations for Reasoning of Unintentional Activities
by: Grover, Shresth, et al.
Published: (2024)
by: Grover, Shresth, et al.
Published: (2024)
On the Limitations of Vision-Language Models in Understanding Image Transforms
by: Anis, Ahmad Mustafa, et al.
Published: (2025)
by: Anis, Ahmad Mustafa, et al.
Published: (2025)
On Occlusions in Video Action Detection: Benchmark Datasets And Training Recipes
by: Modi, Rajat, et al.
Published: (2024)
by: Modi, Rajat, et al.
Published: (2024)
Similar Items
-
Inference-Time Scaling for Complex Tasks: Where We Stand and What Lies Ahead
by: Balachandran, Vidhisha, et al.
Published: (2025) -
Unearthing Skill-Level Insights for Understanding Trade-Offs of Foundation Models
by: Moayeri, Mazda, et al.
Published: (2024) -
MM-GEN: Enhancing Task Performance Through Targeted Multimodal Data Curation
by: Joshi, Siddharth, et al.
Published: (2025) -
Improving Instruction-Following in Language Models through Activation Steering
by: Stolfo, Alessandro, et al.
Published: (2024) -
BenchAgents: Multi-Agent Systems for Structured Benchmark Creation
by: Butt, Natasha, et al.
Published: (2024)