The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
Fuente:
arXiv
Salvato in:
| Autori principali: | Ashury-Tahan, Shir, Mai, Yifan, C, Rajmohan, Gera, Ariel, Perlitz, Yotam, Yehudai, Asaf, Bandel, Elron, Choshen, Leshem, Shnarch, Eyal, Liang, Percy, Shmueli-Scheuer, Michal |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Robustness as an Emergent Property of Task Performance
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2026)
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2026)
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
di: Perlitz, Yotam, et al.
Pubblicazione: (2024)
di: Perlitz, Yotam, et al.
Pubblicazione: (2024)
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2026)
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2026)
Efficient Benchmarking of Language Models
di: Perlitz, Yotam, et al.
Pubblicazione: (2023)
di: Perlitz, Yotam, et al.
Pubblicazione: (2023)
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
di: Habba, Eliya, et al.
Pubblicazione: (2026)
di: Habba, Eliya, et al.
Pubblicazione: (2026)
Label-Efficient Model Selection for Text Generation
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2024)
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2024)
DOVE: A Large-Scale Multi-Dimensional Predictions Dataset Towards Meaningful LLM Evaluation
di: Habba, Eliya, et al.
Pubblicazione: (2025)
di: Habba, Eliya, et al.
Pubblicazione: (2025)
Unitxt: Flexible, Shareable and Reusable Data Preparation and Evaluation for Generative AI
di: Bandel, Elron, et al.
Pubblicazione: (2024)
di: Bandel, Elron, et al.
Pubblicazione: (2024)
CLEAR: Error Analysis via LLM-as-a-Judge Made Easy
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks
di: Keren, Tomer, et al.
Pubblicazione: (2026)
di: Keren, Tomer, et al.
Pubblicazione: (2026)
Guided Query Refinement: Multimodal Hybrid Retrieval with Test-Time Optimization
di: Uzan, Omri, et al.
Pubblicazione: (2025)
di: Uzan, Omri, et al.
Pubblicazione: (2025)
General Agent Evaluation
di: Bandel, Elron, et al.
Pubblicazione: (2026)
di: Bandel, Elron, et al.
Pubblicazione: (2026)
Agentic CLEAR: Automating Multi-Level Evaluation of LLM Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2026)
di: Yehudai, Asaf, et al.
Pubblicazione: (2026)
Task-Adaptive Embedding Refinement via Test-time LLM Guidance
di: Gera, Ariel, et al.
Pubblicazione: (2026)
di: Gera, Ariel, et al.
Pubblicazione: (2026)
JuStRank: Benchmarking LLM Judges for System Ranking
di: Gera, Ariel, et al.
Pubblicazione: (2024)
di: Gera, Ariel, et al.
Pubblicazione: (2024)
Genie: Achieving Human Parity in Content-Grounded Datasets Generation
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
Instructions Shape Production of Language, not Processing
di: Waldis, Andreas, et al.
Pubblicazione: (2026)
di: Waldis, Andreas, et al.
Pubblicazione: (2026)
When LLMs are Unfit Use FastFit: Fast and Effective Text Classification with Many Classes
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
Teaching Values to Machines: Simulating Human-Like Behavior in LLMs
di: Yehudai, Asaf, et al.
Pubblicazione: (2026)
di: Yehudai, Asaf, et al.
Pubblicazione: (2026)
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
di: Waldis, Andreas, et al.
Pubblicazione: (2024)
di: Waldis, Andreas, et al.
Pubblicazione: (2024)
Mediocrity is the key for LLM as a Judge Anchor Selection
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2026)
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2026)
How Safe is Your Safety Metric? Automatic Concatenation Tests for Metric Reliability
di: Fandina, Ora Nova, et al.
Pubblicazione: (2024)
di: Fandina, Ora Nova, et al.
Pubblicazione: (2024)
Will it Merge? On The Causes of Model Mergeability
di: Rahamim, Adir, et al.
Pubblicazione: (2026)
di: Rahamim, Adir, et al.
Pubblicazione: (2026)
WildIFEval: Instruction Following in the Wild
di: Lior, Gili, et al.
Pubblicazione: (2025)
di: Lior, Gili, et al.
Pubblicazione: (2025)
Time to REFLECT: Can We Trust LLM Judges for Evidence-based Research Agents?
di: Wang, Leyao, et al.
Pubblicazione: (2026)
di: Wang, Leyao, et al.
Pubblicazione: (2026)
Survey on Evaluation of LLM-based Agents
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
di: Yehudai, Asaf, et al.
Pubblicazione: (2025)
Data-driven Coreference-based Ontology Building
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2024)
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2024)
CUBE: A Standard for Unifying Agent Benchmarks
di: Lacoste, Alexandre, et al.
Pubblicazione: (2026)
di: Lacoste, Alexandre, et al.
Pubblicazione: (2026)
Can Gradient Descent Simulate Prompting?
di: Zhang, Eric, et al.
Pubblicazione: (2025)
di: Zhang, Eric, et al.
Pubblicazione: (2025)
A Hitchhiker's Guide to Scaling Law Estimation
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
di: Choshen, Leshem, et al.
Pubblicazione: (2024)
Fuse to Forget: Bias Reduction and Selective Memorization through Model Fusion
di: Zaman, Kerem, et al.
Pubblicazione: (2023)
di: Zaman, Kerem, et al.
Pubblicazione: (2023)
Think Again! The Effect of Test-Time Compute on Preferences, Opinions, and Beliefs of Large Language Models
di: Kour, George, et al.
Pubblicazione: (2025)
di: Kour, George, et al.
Pubblicazione: (2025)
The ShareLM Collection and Plugin: Contributing Human-Model Chats for the Benefit of the Community
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2024)
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2024)
Naturally Occurring Feedback is Common, Extractable and Useful
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2024)
di: Don-Yehiya, Shachar, et al.
Pubblicazione: (2024)
A Nurse is Blue and Elephant is Rugby: Cross Domain Alignment in Large Language Models Reveal Human-like Patterns
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
di: Yehudai, Asaf, et al.
Pubblicazione: (2024)
Onion De Bruijn Sequences: Fixed-Window Counting by Growing the Alphabet
di: Genosar, Dor, et al.
Pubblicazione: (2019)
di: Genosar, Dor, et al.
Pubblicazione: (2019)
ComPEFT: Compression for Communicating Parameter Efficient Updates via Sparsification and Quantization
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
di: Yadav, Prateek, et al.
Pubblicazione: (2023)
Mighty Tracker -- Performance Studies of the MightyPix for LHCb
di: Schmitz, Hannah, et al.
Pubblicazione: (2024)
di: Schmitz, Hannah, et al.
Pubblicazione: (2024)
Jump to Conclusions: Short-Cutting Transformers With Linear Transformations
di: Din, Alexander Yom, et al.
Pubblicazione: (2023)
di: Din, Alexander Yom, et al.
Pubblicazione: (2023)
PolySQL: Scaling Text-to-SQL Evaluation Across SQL Dialects via Automated Backend Isomorphism
di: Perlitz, Yotam, et al.
Pubblicazione: (2026)
di: Perlitz, Yotam, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Robustness as an Emergent Property of Task Performance
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2026) -
Do These LLM Benchmarks Agree? Fixing Benchmark Evaluation with BenchBench
di: Perlitz, Yotam, et al.
Pubblicazione: (2024) -
ErrorMap and ErrorAtlas: Charting the Failure Landscape of Large Language Models
di: Ashury-Tahan, Shir, et al.
Pubblicazione: (2026) -
Efficient Benchmarking of Language Models
di: Perlitz, Yotam, et al.
Pubblicazione: (2023) -
Growing Pains: Extensible and Efficient LLM Benchmarking Via Fixed Parameter Calibration
di: Habba, Eliya, et al.
Pubblicazione: (2026)