InfoSynth: Information-Guided Benchmark Synthesis for LLMs
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Garg, Ishir, Kolhe, Neel, Zhao, Xuandong, Song, Dawn |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MemFail: Stress-Testing Failure Modes of LLM Memory Systems
von: Garg, Ishir, et al.
Veröffentlicht: (2026)
von: Garg, Ishir, et al.
Veröffentlicht: (2026)
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
von: Xie, Jingxu, et al.
Veröffentlicht: (2025)
von: Xie, Jingxu, et al.
Veröffentlicht: (2025)
Fisher-Orthogonal Projected Natural Gradient Descent for Continual Learning
von: Garg, Ishir, et al.
Veröffentlicht: (2026)
von: Garg, Ishir, et al.
Veröffentlicht: (2026)
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
von: Xiong, Alexander, et al.
Veröffentlicht: (2025)
von: Xiong, Alexander, et al.
Veröffentlicht: (2025)
Scalable Best-of-N Selection for Large Language Models via Self-Certainty
von: Kang, Zhewei, et al.
Veröffentlicht: (2025)
von: Kang, Zhewei, et al.
Veröffentlicht: (2025)
Are You Getting What You Pay For? Auditing Model Substitution in LLM APIs
von: Cai, Will, et al.
Veröffentlicht: (2025)
von: Cai, Will, et al.
Veröffentlicht: (2025)
In-Context Watermarks for Large Language Models
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
Learning to Reason without External Rewards
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
Position: LLM Watermarking Should Align Stakeholders' Incentives for Practical Adoption
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
von: Liu, Yepeng, et al.
Veröffentlicht: (2025)
CodeHalu: Investigating Code Hallucinations in LLMs via Execution-based Verification
von: Tian, Yuchen, et al.
Veröffentlicht: (2024)
von: Tian, Yuchen, et al.
Veröffentlicht: (2024)
WideSearch: Benchmarking Agentic Broad Info-Seeking
von: Wong, Ryan, et al.
Veröffentlicht: (2025)
von: Wong, Ryan, et al.
Veröffentlicht: (2025)
Multimodal Situational Safety
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2024)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2024)
Improving LLM Safety Alignment with Dual-Objective Optimization
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2025)
Permute-and-Flip: An optimally stable and watermarkable decoder for LLMs
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
Assessing Judging Bias in Large Reasoning Models: An Empirical Study
von: Wang, Qian, et al.
Veröffentlicht: (2025)
von: Wang, Qian, et al.
Veröffentlicht: (2025)
InfoDeepSeek: Benchmarking Agentic Information Seeking for Retrieval-Augmented Generation
von: Xi, Yunjia, et al.
Veröffentlicht: (2025)
von: Xi, Yunjia, et al.
Veröffentlicht: (2025)
Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
von: Liang, Kaiqu, et al.
Veröffentlicht: (2025)
CultureSynth: A Hierarchical Taxonomy-Guided and Retrieval-Augmented Framework for Cultural Question-Answer Synthesis
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
SafeKey: Amplifying Aha-Moment Insights for Safety Reasoning
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
von: Zhou, Kaiwen, et al.
Veröffentlicht: (2025)
MateInfoUB: A Real-World Benchmark for Testing LLMs in Competitive, Multilingual, and Multimodal Educational Tasks
von: Marius, Dumitran Adrian, et al.
Veröffentlicht: (2025)
von: Marius, Dumitran Adrian, et al.
Veröffentlicht: (2025)
SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators
von: Moskovskiy, Daniil, et al.
Veröffentlicht: (2025)
von: Moskovskiy, Daniil, et al.
Veröffentlicht: (2025)
InfoLossQA: Characterizing and Recovering Information Loss in Text Simplification
von: Trienes, Jan, et al.
Veröffentlicht: (2024)
von: Trienes, Jan, et al.
Veröffentlicht: (2024)
A Practical Examination of AI-Generated Text Detectors for Large Language Models
von: Tufts, Brian, et al.
Veröffentlicht: (2024)
von: Tufts, Brian, et al.
Veröffentlicht: (2024)
InfoAgent: Advancing Autonomous Information-Seeking Agents
von: Zhang, Gongrui, et al.
Veröffentlicht: (2025)
von: Zhang, Gongrui, et al.
Veröffentlicht: (2025)
InfoTech Assistant: A Multimodal Conversational Agent for InfoTechnology Web Portal Queries
von: Gadiraju, Sai Surya, et al.
Veröffentlicht: (2024)
von: Gadiraju, Sai Surya, et al.
Veröffentlicht: (2024)
Hidden Persuaders: LLMs' Political Leaning and Their Influence on Voters
von: Potter, Yujin, et al.
Veröffentlicht: (2024)
von: Potter, Yujin, et al.
Veröffentlicht: (2024)
SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis
von: Wu, Zijian, et al.
Veröffentlicht: (2025)
von: Wu, Zijian, et al.
Veröffentlicht: (2025)
InfoGatherer: Principled Information Seeking via Evidence Retrieval and Strategic Questioning
von: Taranukhin, Maksym, et al.
Veröffentlicht: (2026)
von: Taranukhin, Maksym, et al.
Veröffentlicht: (2026)
Reliable Fine-Grained Evaluation of Natural Language Math Proofs
von: Ma, Wenjie, et al.
Veröffentlicht: (2025)
von: Ma, Wenjie, et al.
Veröffentlicht: (2025)
DeepSynth-Eval: Objectively Evaluating Information Consolidation in Deep Survey Writing
von: Zhang, Hongzhi, et al.
Veröffentlicht: (2026)
von: Zhang, Hongzhi, et al.
Veröffentlicht: (2026)
Let's Use ChatGPT To Write Our Paper! Benchmarking LLMs To Write the Introduction of a Research Paper
von: Garg, Krishna, et al.
Veröffentlicht: (2025)
von: Garg, Krishna, et al.
Veröffentlicht: (2025)
InfoDensity: Rewarding Information-Dense Traces for Efficient Reasoning
von: Wei, Chengwei, et al.
Veröffentlicht: (2026)
von: Wei, Chengwei, et al.
Veröffentlicht: (2026)
InfoFlood: Jailbreaking Large Language Models with Information Overload
von: Yadav, Advait, et al.
Veröffentlicht: (2025)
von: Yadav, Advait, et al.
Veröffentlicht: (2025)
PennySynth: RAG-Driven Data Synthesis for Automated Quantum Code Generation
von: Shao, Minghao, et al.
Veröffentlicht: (2026)
von: Shao, Minghao, et al.
Veröffentlicht: (2026)
InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic Modeling
von: Wu, Xiaobao, et al.
Veröffentlicht: (2023)
von: Wu, Xiaobao, et al.
Veröffentlicht: (2023)
Efficiently Identifying Watermarked Segments in Mixed-Source Texts
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
von: Zhao, Xuandong, et al.
Veröffentlicht: (2024)
Why and How LLMs Hallucinate: Connecting the Dots with Subsequence Associations
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
von: Sun, Yiyou, et al.
Veröffentlicht: (2025)
Summing Up the Facts: Additive Mechanisms Behind Factual Recall in LLMs
von: Chughtai, Bilal, et al.
Veröffentlicht: (2024)
von: Chughtai, Bilal, et al.
Veröffentlicht: (2024)
Data Cartography for Detecting Memorization Hotspots and Guiding Data Interventions in Generative Models
von: Patel, Laksh, et al.
Veröffentlicht: (2025)
von: Patel, Laksh, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
MemFail: Stress-Testing Failure Modes of LLM Memory Systems
von: Garg, Ishir, et al.
Veröffentlicht: (2026) -
AgentSynth: Scalable Task Generation for Generalist Computer-Use Agents
von: Xie, Jingxu, et al.
Veröffentlicht: (2025) -
Fisher-Orthogonal Projected Natural Gradient Descent for Continual Learning
von: Garg, Ishir, et al.
Veröffentlicht: (2026) -
Dataset Protection via Watermarked Canaries in Retrieval-Augmented LLMs
von: Liu, Yepeng, et al.
Veröffentlicht: (2025) -
The Landscape of Memorization in LLMs: Mechanisms, Measurement, and Mitigation
von: Xiong, Alexander, et al.
Veröffentlicht: (2025)