SealQA: Raising the Bar for Reasoning in Search-Augmented Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Pham, Thinh, Nguyen, Nguyen, Zunjare, Pratibha, Chen, Weiyuan, Tseng, Yu-Min, Vu, Tu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
$π^2$: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models
von: Do, Quyet V., et al.
Veröffentlicht: (2026)
von: Do, Quyet V., et al.
Veröffentlicht: (2026)
NeuroProlog: Multi-Task Fine-Tuning for Neurosymbolic Mathematical Reasoning via the Cocktail Effect
von: Zunjare, Pratibha, et al.
Veröffentlicht: (2026)
von: Zunjare, Pratibha, et al.
Veröffentlicht: (2026)
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
von: Wang, Han, et al.
Veröffentlicht: (2026)
von: Wang, Han, et al.
Veröffentlicht: (2026)
Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi
von: Fatimah, Shiza, et al.
Veröffentlicht: (2026)
von: Fatimah, Shiza, et al.
Veröffentlicht: (2026)
Formal Reasoning for Intelligent QA Systems: A Case Study in the Educational Domain
von: Bui, Tuan, et al.
Veröffentlicht: (2025)
von: Bui, Tuan, et al.
Veröffentlicht: (2025)
VLegal-Bench: Cognitively Grounded Benchmark for Vietnamese Legal Reasoning of Large Language Models
von: Dong, Nguyen Tien, et al.
Veröffentlicht: (2025)
von: Dong, Nguyen Tien, et al.
Veröffentlicht: (2025)
Tougher Text, Smarter Models: Raising the Bar for Adversarial Defence Benchmarks
von: Wang, Yang, et al.
Veröffentlicht: (2025)
von: Wang, Yang, et al.
Veröffentlicht: (2025)
Raising the Bar: Investigating the Values of Large Language Models via Generative Evolving Testing
von: Jiang, Han, et al.
Veröffentlicht: (2024)
von: Jiang, Han, et al.
Veröffentlicht: (2024)
Reasoning Planning for Language Models
von: Nguyen, Bao, et al.
Veröffentlicht: (2025)
von: Nguyen, Bao, et al.
Veröffentlicht: (2025)
GPTs and Language Barrier: A Cross-Lingual Legal QA Examination
von: Nguyen, Ha-Thanh, et al.
Veröffentlicht: (2024)
von: Nguyen, Ha-Thanh, et al.
Veröffentlicht: (2024)
Leveraging Large Language Models for Suicide Detection on Social Media with Limited Labels
von: Nguyen, Vy, et al.
Veröffentlicht: (2024)
von: Nguyen, Vy, et al.
Veröffentlicht: (2024)
MedBioRAG: Semantic Search and Retrieval-Augmented Generation with Large Language Models for Medical and Biological QA
von: Kim, Seonok
Veröffentlicht: (2025)
von: Kim, Seonok
Veröffentlicht: (2025)
A Hybrid Multi-Agent Prompting Approach for Simplifying Complex Sentences
von: Zunjare, Pratibha, et al.
Veröffentlicht: (2025)
von: Zunjare, Pratibha, et al.
Veröffentlicht: (2025)
MA-RAG: Multi-Agent Retrieval-Augmented Generation via Collaborative Chain-of-Thought Reasoning
von: Nguyen, Thang, et al.
Veröffentlicht: (2025)
von: Nguyen, Thang, et al.
Veröffentlicht: (2025)
Bridging the Reasoning Gap in Vietnamese with Small Language Models via Test-Time Scaling
von: Trung, Bui The, et al.
Veröffentlicht: (2026)
von: Trung, Bui The, et al.
Veröffentlicht: (2026)
BERT-based model for Vietnamese Fact Verification Dataset
von: Tran, Bao, et al.
Veröffentlicht: (2025)
von: Tran, Bao, et al.
Veröffentlicht: (2025)
PaperSearchQA: Learning to Search and Reason over Scientific Papers with RLVR
von: Burgess, James, et al.
Veröffentlicht: (2026)
von: Burgess, James, et al.
Veröffentlicht: (2026)
Enhancing Retrieval Augmented Generation with Hierarchical Text Segmentation Chunking
von: Nguyen, Hai Toan, et al.
Veröffentlicht: (2025)
von: Nguyen, Hai Toan, et al.
Veröffentlicht: (2025)
Few-shot Continual Relation Extraction via Open Information Extraction
von: Nguyen, Thiem, et al.
Veröffentlicht: (2025)
von: Nguyen, Thiem, et al.
Veröffentlicht: (2025)
LIBMoE: A Library for comprehensive benchmarking Mixture of Experts in Large Language Models
von: Nguyen, Nam V., et al.
Veröffentlicht: (2024)
von: Nguyen, Nam V., et al.
Veröffentlicht: (2024)
Where Knowledge Collides: A Mechanistic Study of Intra-Memory Knowledge Conflict in Language Models
von: Pham, Minh Vu, et al.
Veröffentlicht: (2026)
von: Pham, Minh Vu, et al.
Veröffentlicht: (2026)
SuperRAG: Beyond RAG with Layout-Aware Graph Modeling
von: Yang, Jeff, et al.
Veröffentlicht: (2025)
von: Yang, Jeff, et al.
Veröffentlicht: (2025)
PRISM: Pushing the Frontier of Deep Think via Process Reward Model-Guided Inference
von: Sharma, Rituraj, et al.
Veröffentlicht: (2026)
von: Sharma, Rituraj, et al.
Veröffentlicht: (2026)
Bridging LLMs and Symbolic Reasoning in Educational QA Systems: Insights from the XAI Challenge at IJCNN 2025
von: Nguyen, Long S. T., et al.
Veröffentlicht: (2025)
von: Nguyen, Long S. T., et al.
Veröffentlicht: (2025)
Robust Search with Uncertainty-Aware Value Models for Language Model Reasoning
von: Yu, Fei, et al.
Veröffentlicht: (2025)
von: Yu, Fei, et al.
Veröffentlicht: (2025)
Tree-OPO: Off-policy Monte Carlo Tree-Guided Advantage Optimization for Multistep Reasoning
von: Huang, Bingning, et al.
Veröffentlicht: (2025)
von: Huang, Bingning, et al.
Veröffentlicht: (2025)
Improving Retrieval Augmented Language Model with Self-Reasoning
von: Xia, Yuan, et al.
Veröffentlicht: (2024)
von: Xia, Yuan, et al.
Veröffentlicht: (2024)
OpenSeal: Good, Fast, and Cheap Construction of an Open-Source Southeast Asian LLM via Parallel Data
von: Nguyen, Tan Sang, et al.
Veröffentlicht: (2026)
von: Nguyen, Tan Sang, et al.
Veröffentlicht: (2026)
Speaking in Words, Thinking in Logic: A Dual-Process Framework in QA Systems
von: Bui, Tuan, et al.
Veröffentlicht: (2025)
von: Bui, Tuan, et al.
Veröffentlicht: (2025)
HiQA: A Hierarchical Contextual Augmentation RAG for Multi-Documents QA
von: Chen, Xinyue, et al.
Veröffentlicht: (2024)
von: Chen, Xinyue, et al.
Veröffentlicht: (2024)
Multilingual Text-to-SQL: Benchmarking the Limits of Language Models with Collaborative Language Agents
von: Pham, Khanh Trinh, et al.
Veröffentlicht: (2025)
von: Pham, Khanh Trinh, et al.
Veröffentlicht: (2025)
InfoCausalQA:Can Models Perform Non-explicit Causal Reasoning Based on Infographic?
von: Ka, Keummin, et al.
Veröffentlicht: (2025)
von: Ka, Keummin, et al.
Veröffentlicht: (2025)
Retrospex: Language Agent Meets Offline Reinforcement Learning Critic
von: Xiang, Yufei, et al.
Veröffentlicht: (2025)
von: Xiang, Yufei, et al.
Veröffentlicht: (2025)
Learning to Compress: Unlocking the Potential of Large Language Models for Text Representation
von: Zhang, Yeqin, et al.
Veröffentlicht: (2025)
von: Zhang, Yeqin, et al.
Veröffentlicht: (2025)
CompeteSMoE -- Statistically Guaranteed Mixture of Experts Training via Competition
von: Nguyen, Nam V., et al.
Veröffentlicht: (2025)
von: Nguyen, Nam V., et al.
Veröffentlicht: (2025)
Discourse Graph Guided Document Translation with Large Language Models
von: Pham, Viet-Thanh, et al.
Veröffentlicht: (2025)
von: Pham, Viet-Thanh, et al.
Veröffentlicht: (2025)
Enhancing Legal Document Retrieval: A Multi-Phase Approach with Large Language Models
von: Nguyen, Hai-Long, et al.
Veröffentlicht: (2024)
von: Nguyen, Hai-Long, et al.
Veröffentlicht: (2024)
Intra-Layer Recurrence in Transformers for Language Modeling
von: Nguyen, Anthony, et al.
Veröffentlicht: (2025)
von: Nguyen, Anthony, et al.
Veröffentlicht: (2025)
ARise: Towards Knowledge-Augmented Reasoning via Risk-Adaptive Search
von: Zhang, Yize, et al.
Veröffentlicht: (2025)
von: Zhang, Yize, et al.
Veröffentlicht: (2025)
MedBioLM: Optimizing Medical and Biological QA with Fine-Tuned Large Language Models and Retrieval-Augmented Generation
von: Kim, Seonok
Veröffentlicht: (2025)
von: Kim, Seonok
Veröffentlicht: (2025)
Ähnliche Einträge
-
$π^2$: Structure-Originated Reasoning Data Improves Long-Context Reasoning Ability of Large Language Models
von: Do, Quyet V., et al.
Veröffentlicht: (2026) -
NeuroProlog: Multi-Task Fine-Tuning for Neurosymbolic Mathematical Reasoning via the Cocktail Effect
von: Zunjare, Pratibha, et al.
Veröffentlicht: (2026) -
MERRIN: A Benchmark for Multimodal Evidence Retrieval and Reasoning in Noisy Web Environments
von: Wang, Han, et al.
Veröffentlicht: (2026) -
Raising Bars, Not Parameters: LilMoo Compact Language Model for Hindi
von: Fatimah, Shiza, et al.
Veröffentlicht: (2026) -
Formal Reasoning for Intelligent QA Systems: A Case Study in the Educational Domain
von: Bui, Tuan, et al.
Veröffentlicht: (2025)