Overview of the Sensemaking Task at the ELOQUENT 2025 Lab: LLMs as Teachers, Students and Evaluators
Fuente:
arXiv
Saved in:
| Main Authors: | Šindelář, Pavel, Bojar, Ondřej |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
by: Jiang, Yilin, et al.
Published: (2025)
by: Jiang, Yilin, et al.
Published: (2025)
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025)
by: Saji, Alan, et al.
Published: (2025)
LLMs Are Not Scorers: Rethinking MT Evaluation with Generation-Based Methods
by: Cui, Hyang
Published: (2025)
by: Cui, Hyang
Published: (2025)
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)
by: Ashuach, Tomer, et al.
Published: (2025)
Overview of the ClinIQLink 2025 Shared Task on Medical Question-Answering
by: Colelough, Brandon, et al.
Published: (2025)
by: Colelough, Brandon, et al.
Published: (2025)
LombardoGraphia: Automatic Classification of Lombard Orthography Variants
by: Signoroni, Edoardo, et al.
Published: (2026)
by: Signoroni, Edoardo, et al.
Published: (2026)
OpenFactCheck: Building, Benchmarking Customized Fact-Checking Systems and Evaluating the Factuality of Claims and LLMs
by: Wang, Yuxia, et al.
Published: (2024)
by: Wang, Yuxia, et al.
Published: (2024)
The CLEF-2025 CheckThat! Lab: Subjectivity, Fact-Checking, Claim Normalization, and Retrieval
by: Alam, Firoj, et al.
Published: (2025)
by: Alam, Firoj, et al.
Published: (2025)
Automatic Task Detection and Heterogeneous LLM Speculative Decoding
by: Ge, Danying, et al.
Published: (2025)
by: Ge, Danying, et al.
Published: (2025)
Constructing Benchmarks and Interventions for Combating Hallucinations in LLMs
by: Simhi, Adi, et al.
Published: (2024)
by: Simhi, Adi, et al.
Published: (2024)
University of Indonesia at SemEval-2025 Task 11: Evaluating State-of-the-Art Encoders for Multi-Label Emotion Detection
by: Hanif, Ikhlasul Akmal, et al.
Published: (2025)
by: Hanif, Ikhlasul Akmal, et al.
Published: (2025)
Task Contamination: Language Models May Not Be Few-Shot Anymore
by: Li, Changmao, et al.
Published: (2023)
by: Li, Changmao, et al.
Published: (2023)
Adapting Multilingual Models to Code-Mixed Tasks via Model Merging
by: Kodali, Prashant, et al.
Published: (2025)
by: Kodali, Prashant, et al.
Published: (2025)
Neither Valid nor Reliable? Investigating the Use of LLMs as Judges
by: Chehbouni, Khaoula, et al.
Published: (2025)
by: Chehbouni, Khaoula, et al.
Published: (2025)
MIRIAD: Augmenting LLMs with millions of medical query-response pairs
by: Zheng, Qinyue, et al.
Published: (2025)
by: Zheng, Qinyue, et al.
Published: (2025)
RAG-Optimized Tibetan Tourism LLMs: Enhancing Accuracy and Personalization
by: Qi, Jinhu, et al.
Published: (2024)
by: Qi, Jinhu, et al.
Published: (2024)
MAWARITH: A Dataset and Benchmark for Legal Inheritance Reasoning with LLMs
by: Bouchekif, Abdessalam, et al.
Published: (2026)
by: Bouchekif, Abdessalam, et al.
Published: (2026)
HACHIMI: Scalable and Controllable Student Persona Generation via Orchestrated Agents
by: Jiang, Yilin, et al.
Published: (2026)
by: Jiang, Yilin, et al.
Published: (2026)
LLMs and the Human Condition
by: Wallis, Peter
Published: (2024)
by: Wallis, Peter
Published: (2024)
A Multi-Task Benchmark for Abusive Language Detection in Low-Resource Settings
by: Gaim, Fitsum, et al.
Published: (2025)
by: Gaim, Fitsum, et al.
Published: (2025)
Fine-Tuning LLMs with Fine-Grained Human Feedback on Text Spans
by: CH-Wang, Sky, et al.
Published: (2025)
by: CH-Wang, Sky, et al.
Published: (2025)
Fane at SemEval-2025 Task 10: Zero-Shot Entity Framing with Large Language Models
by: Fane, Enfa, et al.
Published: (2025)
by: Fane, Enfa, et al.
Published: (2025)
Just Pass Twice: Efficient Token Classification with LLMs for Zero-Shot NER
by: Ewais, Ahmed, et al.
Published: (2026)
by: Ewais, Ahmed, et al.
Published: (2026)
KGiRAG: An Iterative GraphRAG Approach for Responding Sensemaking Queries
by: Iacob, Isabela, et al.
Published: (2026)
by: Iacob, Isabela, et al.
Published: (2026)
GroUSE: A Benchmark to Evaluate Evaluators in Grounded Question Answering
by: Muller, Sacha, et al.
Published: (2024)
by: Muller, Sacha, et al.
Published: (2024)
SemEval-2026 Task 3: Dimensional Aspect-Based Sentiment Analysis (DimABSA)
by: Yu, Liang-Chih, et al.
Published: (2026)
by: Yu, Liang-Chih, et al.
Published: (2026)
Language Models Can Resolve Reference Compositionally, But It's Not Their Native Strength: The Case of the Personal Relation Task
by: Evelo, Bart, et al.
Published: (2026)
by: Evelo, Bart, et al.
Published: (2026)
MALT: Mechanistic Ablation of Lossy Translation in LLMs for a Low-Resource Language: Urdu
by: Bajwa, Taaha Saleem
Published: (2025)
by: Bajwa, Taaha Saleem
Published: (2025)
Trust Me, I'm Wrong: LLMs Hallucinate with Certainty Despite Knowing the Answer
by: Simhi, Adi, et al.
Published: (2025)
by: Simhi, Adi, et al.
Published: (2025)
Temporal Referential Consistency: Do LLMs Favor Sequences Over Absolute Time References?
by: Bajpai, Ashutosh, et al.
Published: (2025)
by: Bajpai, Ashutosh, et al.
Published: (2025)
Improving the OOD Performance of Closed-Source LLMs on NLI Through Strategic Data Selection
by: Stacey, Joe, et al.
Published: (2025)
by: Stacey, Joe, et al.
Published: (2025)
A Case Study of Cross-Lingual Zero-Shot Generalization for Classical Languages in LLMs
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2025)
by: Akavarapu, V. S. D. S. Mahesh, et al.
Published: (2025)
Benchmarking the Performance of Pre-trained LLMs across Urdu NLP Tasks
by: Tahir, Munief Hassan, et al.
Published: (2024)
by: Tahir, Munief Hassan, et al.
Published: (2024)
Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review
by: Peters, Sydney, et al.
Published: (2025)
by: Peters, Sydney, et al.
Published: (2025)
Homogeneous Keys, Heterogeneous Values: Exploiting Local KV Cache Asymmetry for Long-Context LLMs
by: Cui, Wanyun, et al.
Published: (2025)
by: Cui, Wanyun, et al.
Published: (2025)
CascadeMind at SemEval-2026 Task 4: A Hybrid Neuro-Symbolic Cascade for Narrative Similarity
by: Kawada, Sebastien, et al.
Published: (2026)
by: Kawada, Sebastien, et al.
Published: (2026)
All for One: LLMs Solve Mental Math at the Last Token With Information Transferred From Other Tokens
by: Mamidanna, Siddarth, et al.
Published: (2025)
by: Mamidanna, Siddarth, et al.
Published: (2025)
Heidelberg-Boston @ SIGTYP 2024 Shared Task: Enhancing Low-Resource Language Analysis With Character-Aware Hierarchical Transformers
by: Riemenschneider, Frederick, et al.
Published: (2024)
by: Riemenschneider, Frederick, et al.
Published: (2024)
Improving LLMs with a knowledge from databases
by: Máša, Petr
Published: (2025)
by: Máša, Petr
Published: (2025)
Similar Items
-
EduGuardBench: A Holistic Benchmark for Evaluating the Pedagogical Fidelity and Adversarial Safety of LLMs as Simulated Teachers
by: Jiang, Yilin, et al.
Published: (2025) -
RomanLens: The Role Of Latent Romanization In Multilinguality In LLMs
by: Saji, Alan, et al.
Published: (2025) -
LLMs Are Not Scorers: Rethinking MT Evaluation with Generation-Based Methods
by: Cui, Hyang
Published: (2025) -
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
by: Simhi, Adi, et al.
Published: (2025) -
CRISP: Persistent Concept Unlearning via Sparse Autoencoders
by: Ashuach, Tomer, et al.
Published: (2025)