Tiny QA Benchmark++: Ultra-Lightweight, Synthetic Multilingual Dataset Generation & Smoke-Tests for Continuous LLM Evaluation
Fuente:
arXiv
Saved in:
| Main Author: | Koc, Vincent |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Survey Transfer Learning: Recycling Data with Silicon Responses
by: Amini, Ali
Published: (2025)
by: Amini, Ali
Published: (2025)
Evo-DKD: Dual-Knowledge Decoding for Autonomous Ontology Evolution in Large Language Models
by: Raman, Vishal, et al.
Published: (2025)
by: Raman, Vishal, et al.
Published: (2025)
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
by: Zhou, Yue, et al.
Published: (2025)
by: Zhou, Yue, et al.
Published: (2025)
GraphEval36K: Benchmarking Coding and Reasoning Capabilities of Large Language Models on Graph Datasets
by: Wu, Qiming, et al.
Published: (2024)
by: Wu, Qiming, et al.
Published: (2024)
ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs
by: Cui, Jian, et al.
Published: (2026)
by: Cui, Jian, et al.
Published: (2026)
The Single-File Test: A Longitudinal Public-Interface Evaluation of First-Output LLM Web Generation with Social Reach Tracking
by: Palacios, Diego Cabezas
Published: (2026)
by: Palacios, Diego Cabezas
Published: (2026)
Enhancing Document AI Data Generation Through Graph-Based Synthetic Layouts
by: Agarwal, Amit, et al.
Published: (2024)
by: Agarwal, Amit, et al.
Published: (2024)
Dynamics of COVID-19 Misinformation: An Analysis of Conspiracy Theories, Fake Remedies, and False Reports
by: Thakur, Nirmalya, et al.
Published: (2025)
by: Thakur, Nirmalya, et al.
Published: (2025)
Learning, Fast and Slow: Towards LLMs That Adapt Continually
by: Tiwari, Rishabh, et al.
Published: (2026)
by: Tiwari, Rishabh, et al.
Published: (2026)
SCULPT: Constraint-Guided Pruned MCTS that Carves Efficient Paths for Mathematical Reasoning
by: Fang, Qitong, et al.
Published: (2026)
by: Fang, Qitong, et al.
Published: (2026)
On the Limits of Learned Importance Scoring for KV Cache Compression
by: Steele, Brady
Published: (2026)
by: Steele, Brady
Published: (2026)
Scaling Trends for Multi-Hop Contextual Reasoning in Mid-Scale Language Models
by: Steele, Brady, et al.
Published: (2026)
by: Steele, Brady, et al.
Published: (2026)
The Reasoning-Creativity Trade-off: Toward Creativity-Driven Problem Solving
by: Luyten, Max Ruiz, et al.
Published: (2026)
by: Luyten, Max Ruiz, et al.
Published: (2026)
Dynamic Policy Induction for Adaptive Prompt Optimization: Bridging the Efficiency-Accuracy Gap via Lightweight Reinforcement Learning
by: Xu, Jiexi
Published: (2025)
by: Xu, Jiexi
Published: (2025)
HySemRAG: A Hybrid Semantic Retrieval-Augmented Generation Framework for Automated Literature Synthesis and Methodological Gap Analysis
by: Godinez, Alejandro
Published: (2025)
by: Godinez, Alejandro
Published: (2025)
Streaming Continual Learning for Unified Adaptive Intelligence in Dynamic Environments
by: Giannini, Federico, et al.
Published: (2026)
by: Giannini, Federico, et al.
Published: (2026)
Mpox Narrative on Instagram: A Labeled Multilingual Dataset of Instagram Posts on Mpox for Sentiment, Hate Speech, and Anxiety Analysis
by: Thakur, Nirmalya
Published: (2024)
by: Thakur, Nirmalya
Published: (2024)
Five Years of COVID-19 Discourse on Instagram: A Labeled Instagram Dataset of Over Half a Million Posts for Multilingual Sentiment Analysis
by: Thakur, Nirmalya
Published: (2024)
by: Thakur, Nirmalya
Published: (2024)
Systematic Classification of Studies Investigating Social Media Conversations about Long COVID Using a Novel Zero-Shot Transformer Framework
by: Thakur, Nirmalya, et al.
Published: (2025)
by: Thakur, Nirmalya, et al.
Published: (2025)
Learning Natural Language Constraints for Safe Reinforcement Learning of Language Agents
by: Chua, Jaymari, et al.
Published: (2025)
by: Chua, Jaymari, et al.
Published: (2025)
Vis-CoT: A Human-in-the-Loop Framework for Interactive Visualization and Intervention in LLM Chain-of-Thought Reasoning
by: Pather, Kaviraj, et al.
Published: (2025)
by: Pather, Kaviraj, et al.
Published: (2025)
AI and Machine Learning Approaches for Predicting Nanoparticles Toxicity The Critical Role of Physiochemical Properties
by: Yousaf, Iqra
Published: (2024)
by: Yousaf, Iqra
Published: (2024)
HELEA: Hard-Negative Benchmark and LLM-based Reranking for Robust Entity Alignment
by: Jang, Yoonjin, et al.
Published: (2026)
by: Jang, Yoonjin, et al.
Published: (2026)
Primary Care Diagnoses as a Reliable Predictor for Orthopedic Surgical Interventions
by: Verma, Khushboo, et al.
Published: (2025)
by: Verma, Khushboo, et al.
Published: (2025)
Resolving Action Bottleneck: Agentic Reinforcement Learning Informed by Token-Level Energy
by: He, Langzhou, et al.
Published: (2026)
by: He, Langzhou, et al.
Published: (2026)
Aligning LLMs on a Budget: Inference-Time Alignment with Heuristic Reward Models
by: Nakamura, Mason, et al.
Published: (2025)
by: Nakamura, Mason, et al.
Published: (2025)
XAutoLM: Efficient Fine-Tuning of Language Models via Meta-Learning and AutoML
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
by: Estevanell-Valladares, Ernesto L., et al.
Published: (2025)
A Labelled Dataset for Sentiment Analysis of Videos on YouTube, TikTok, and Other Sources about the 2024 Outbreak of Measles
by: Thakur, Nirmalya, et al.
Published: (2024)
by: Thakur, Nirmalya, et al.
Published: (2024)
Are We Winning the Wrong Game? Revisiting Evaluation Practices for Long-Term Time Series Forecasting
by: Phungtua-eng, Thanapol, et al.
Published: (2026)
by: Phungtua-eng, Thanapol, et al.
Published: (2026)
Predictive Analytics for Collaborators Answers, Code Quality, and Dropout on Stack Overflow
by: Zolduoarrati, Elijah, et al.
Published: (2025)
by: Zolduoarrati, Elijah, et al.
Published: (2025)
MOCHA: Multi-Objective Chebyshev Annealing for Agent Skill Optimization
by: Tanjim, Md Mehrab, et al.
Published: (2026)
by: Tanjim, Md Mehrab, et al.
Published: (2026)
Retrieval Augmented Thought Process for Private Data Handling in Healthcare
by: Pouplin, Thomas, et al.
Published: (2024)
by: Pouplin, Thomas, et al.
Published: (2024)
LLM-as-a-Judge: Rapid Evaluation of Legal Document Recommendation for Retrieval-Augmented Generation
by: Pradhan, Anu, et al.
Published: (2025)
by: Pradhan, Anu, et al.
Published: (2025)
PilotBench: A Benchmark for General Aviation Agents with Safety Constraints
by: Wu, Yalun, et al.
Published: (2026)
by: Wu, Yalun, et al.
Published: (2026)
IMDMR: An Intelligent Multi-Dimensional Memory Retrieval System for Enhanced Conversational AI
by: Pawar, Tejas, et al.
Published: (2025)
by: Pawar, Tejas, et al.
Published: (2025)
COVID-19 on YouTube: A Data-Driven Analysis of Sentiment, Toxicity, and Content Recommendations
by: Su, Vanessa, et al.
Published: (2024)
by: Su, Vanessa, et al.
Published: (2024)
Content and Engagement Trends in COVID-19 YouTube Videos: Evidence from the Late Pandemic
by: Thakur, Nirmalya, et al.
Published: (2025)
by: Thakur, Nirmalya, et al.
Published: (2025)
Emoji Retrieval from Gibberish or Garbled Social Media Text: A Novel Methodology and A Case Study
by: Cui, Shuqi, et al.
Published: (2024)
by: Cui, Shuqi, et al.
Published: (2024)
Quantifying Public Response to COVID-19 Events: Introducing the Community Sentiment and Engagement Index
by: Thakur, Nirmalya, et al.
Published: (2024)
by: Thakur, Nirmalya, et al.
Published: (2024)
Widening the Role of Group Recommender Systems with CAJO
by: Ricci, Francesco, et al.
Published: (2025)
by: Ricci, Francesco, et al.
Published: (2025)
Similar Items
-
Survey Transfer Learning: Recycling Data with Silicon Responses
by: Amini, Ali
Published: (2025) -
Evo-DKD: Dual-Knowledge Decoding for Autonomous Ontology Evolution in Large Language Models
by: Raman, Vishal, et al.
Published: (2025) -
ConfProBench: A Confidence Evaluation Benchmark for MLLM-Based Process Judges
by: Zhou, Yue, et al.
Published: (2025) -
GraphEval36K: Benchmarking Coding and Reasoning Capabilities of Large Language Models on Graph Datasets
by: Wu, Qiming, et al.
Published: (2024) -
ReaGeo: Reasoning-Enhanced End-to-End Geocoding with LLMs
by: Cui, Jian, et al.
Published: (2026)