SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Pham, Thinh, Nguyen, Nguyen, Zunjare, Pratibha, Chen, Weiyuan, Tseng, Yu-Min, Vu, Tu |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
$π^2$: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models
par: Do, Quyet V., et autres
Publié: (2026)
par: Do, Quyet V., et autres
Publié: (2026)
NeuroProlog: Multi-Task Fine-Tuning for Neurosymbolic Mathematical Reasoning via the Cocktail Effect
par: Zunjare, Pratibha, et autres
Publié: (2026)
par: Zunjare, Pratibha, et autres
Publié: (2026)
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
par: Wang, Han, et autres
Publié: (2026)
par: Wang, Han, et autres
Publié: (2026)
Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi
par: Fatimah, Shiza, et autres
Publié: (2026)
par: Fatimah, Shiza, et autres
Publié: (2026)
Formal Reasoning for Intelligent QA Systems: A Case Study in the Educational Domain
par: Bui, Tuan, et autres
Publié: (2025)
par: Bui, Tuan, et autres
Publié: (2025)
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models
par: Dong, Nguyen Tien, et autres
Publié: (2025)
par: Dong, Nguyen Tien, et autres
Publié: (2025)
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks
par: Wang, Yang, et autres
Publié: (2025)
par: Wang, Yang, et autres
Publié: (2025)
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
par: Jiang, Han, et autres
Publié: (2024)
par: Jiang, Han, et autres
Publié: (2024)
Reasoning Planning for Language Models
par: Nguyen, Bao, et autres
Publié: (2025)
par: Nguyen, Bao, et autres
Publié: (2025)
GPTs and Language Barrier: A Cross-Lingual Legal QA Examination
par: Nguyen, Ha-Thanh, et autres
Publié: (2024)
par: Nguyen, Ha-Thanh, et autres
Publié: (2024)
Leveraging Large Language Models for Suicide Detection on Social Media with Limited Labels
par: Nguyen, Vy, et autres
Publié: (2024)
par: Nguyen, Vy, et autres
Publié: (2024)
MedBioRAG: Semantic Search and Retrieval-Augmented Generation with Large Language Models for Medical and Biological QA
par: Kim, Seonok
Publié: (2025)
par: Kim, Seonok
Publié: (2025)
A Hybrid Multi-Agent Prompting Approach for Simplifying Complex Sentences
par: Zunjare, Pratibha, et autres
Publié: (2025)
par: Zunjare, Pratibha, et autres
Publié: (2025)
MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
par: Nguyen, Thang, et autres
Publié: (2025)
par: Nguyen, Thang, et autres
Publié: (2025)
Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling
par: Trung, Bui The, et autres
Publié: (2026)
par: Trung, Bui The, et autres
Publié: (2026)
BERT-based model for Vietnamese Fact Verification Dataset
par: Tran, Bao, et autres
Publié: (2025)
par: Tran, Bao, et autres
Publié: (2025)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
par: Burgess, James, et autres
Publié: (2026)
par: Burgess, James, et autres
Publié: (2026)
Enhancing Retrieval Augmented Generation with Hierarchical Text Segmentation Chunking
par: Nguyen, Hai Toan, et autres
Publié: (2025)
par: Nguyen, Hai Toan, et autres
Publié: (2025)
Few-shot Continual Relation Extraction via Open Information Extraction
par: Nguyen, Thiem, et autres
Publié: (2025)
par: Nguyen, Thiem, et autres
Publié: (2025)
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
par: Nguyen, Nam V., et autres
Publié: (2024)
par: Nguyen, Nam V., et autres
Publié: (2024)
Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models
par: Pham, Minh Vu, et autres
Publié: (2026)
par: Pham, Minh Vu, et autres
Publié: (2026)
SuperRAG: Beyond RAG with Layout-Aware Graph Modeling
par: Yang, Jeff, et autres
Publié: (2025)
par: Yang, Jeff, et autres
Publié: (2025)
PRISM: Pushing the Frontier of Deep Think via Process Reward Model-Guided Inference
par: Sharma, Rituraj, et autres
Publié: (2026)
par: Sharma, Rituraj, et autres
Publié: (2026)
Bridging LLMs and Symbolic Reasoning in Educational QA Systems: Insights from the XAI Challenge at IJCNN 2025
par: Nguyen, Long S. T., et autres
Publié: (2025)
par: Nguyen, Long S. T., et autres
Publié: (2025)
Robust Search with Uncertainty-Aware Value Models for Language Model Reasoning
par: Yu, Fei, et autres
Publié: (2025)
par: Yu, Fei, et autres
Publié: (2025)
Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
par: Huang, Bingning, et autres
Publié: (2025)
par: Huang, Bingning, et autres
Publié: (2025)
Improving Retrieval Augmented Language Model with Self-Reasoning
par: Xia, Yuan, et autres
Publié: (2024)
par: Xia, Yuan, et autres
Publié: (2024)
OpenSeal: Good, Fast, and Cheap Construction of an Open-Source Southeast Asian LLM via Parallel Data
par: Nguyen, Tan Sang, et autres
Publié: (2026)
par: Nguyen, Tan Sang, et autres
Publié: (2026)
Speaking in Words, Thinking in Logic: A Dual-Process Framework in QA Systems
par: Bui, Tuan, et autres
Publié: (2025)
par: Bui, Tuan, et autres
Publié: (2025)
HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA
par: Chen, Xinyue, et autres
Publié: (2024)
par: Chen, Xinyue, et autres
Publié: (2024)
Multilingual Text-to-SQL: Benchmarking the Limits of Language Models with Collaborative Language Agents
par: Pham, Khanh Trinh, et autres
Publié: (2025)
par: Pham, Khanh Trinh, et autres
Publié: (2025)
InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?
par: Ka, Keummin, et autres
Publié: (2025)
par: Ka, Keummin, et autres
Publié: (2025)
Retrospex: Language Agent Meets Offline Reinforcement Learning Critic
par: Xiang, Yufei, et autres
Publié: (2025)
par: Xiang, Yufei, et autres
Publié: (2025)
Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation
par: Zhang, Yeqin, et autres
Publié: (2025)
par: Zhang, Yeqin, et autres
Publié: (2025)
CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition
par: Nguyen, Nam V., et autres
Publié: (2025)
par: Nguyen, Nam V., et autres
Publié: (2025)
Discourse Graph Guided Document Translation with Large Language Models
par: Pham, Viet-Thanh, et autres
Publié: (2025)
par: Pham, Viet-Thanh, et autres
Publié: (2025)
Enhancing Legal Document Retrieval: A Multi-Phase Approach with Large Language Models
par: Nguyen, Hai-Long, et autres
Publié: (2024)
par: Nguyen, Hai-Long, et autres
Publié: (2024)
Intra-Layer Recurrence in Transformers for Language Modeling
par: Nguyen, Anthony, et autres
Publié: (2025)
par: Nguyen, Anthony, et autres
Publié: (2025)
ARise: Towards Knowledge-Augmented Reasoning via Risk-Adaptive Search
par: Zhang, Yize, et autres
Publié: (2025)
par: Zhang, Yize, et autres
Publié: (2025)
MedBioLM: Optimizing Medical and Biological QA with Fine-Tuned Large Language Models and Retrieval-Augmented Generation
par: Kim, Seonok
Publié: (2025)
par: Kim, Seonok
Publié: (2025)
Documents similaires
-
$π^2$: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models
par: Do, Quyet V., et autres
Publié: (2026) -
NeuroProlog: Multi-Task Fine-Tuning for Neurosymbolic Mathematical Reasoning via the Cocktail Effect
par: Zunjare, Pratibha, et autres
Publié: (2026) -
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
par: Wang, Han, et autres
Publié: (2026) -
Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi
par: Fatimah, Shiza, et autres
Publié: (2026) -
Formal Reasoning for Intelligent QA Systems: A Case Study in the Educational Domain
par: Bui, Tuan, et autres
Publié: (2025)