Characterizing Deep Research: A Benchmark and Formal Definition
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Java, Abhinav, Khandelwal, Ashmit, Midigeshi, Sukruta, Halfaker, Aaron, Deshpande, Amit, Goyal, Navin, Gupta, Ankur, Natarajan, Nagarajan, Sharma, Amit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FrugalRAG: Less is More in RL Finetuning for Multi-Hop Question Answering
von: Java, Abhinav, et al.
Veröffentlicht: (2025)
von: Java, Abhinav, et al.
Veröffentlicht: (2025)
Plan*RAG: Efficient Test-Time Planning for Retrieval Augmented Generation
von: Verma, Prakhar, et al.
Veröffentlicht: (2024)
von: Verma, Prakhar, et al.
Veröffentlicht: (2024)
Achieving Limited Adaptivity for Multinomial Logistic Bandits
von: Midigeshi, Sukruta Prakash, et al.
Veröffentlicht: (2025)
von: Midigeshi, Sukruta Prakash, et al.
Veröffentlicht: (2025)
NICE: To Optimize In-Context Examples or Not?
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024)
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024)
Task Facet Learning: A Structured Approach to Prompt Optimization
von: Juneja, Gurusha, et al.
Veröffentlicht: (2024)
von: Juneja, Gurusha, et al.
Veröffentlicht: (2024)
interwhen: A Generalizable Framework for Steering Reasoning Models with Test-time Verification
von: Bhat, Vishak K, et al.
Veröffentlicht: (2026)
von: Bhat, Vishak K, et al.
Veröffentlicht: (2026)
Towards Operationalizing Right to Data Protection
von: Java, Abhinav, et al.
Veröffentlicht: (2024)
von: Java, Abhinav, et al.
Veröffentlicht: (2024)
Counter Turing Test ($CT^2$): Investigating AI-Generated Text Detection for Hindi -- Ranking LLMs based on Hindi AI Detectability Index ($ADI_{hi}$)
von: Kavathekar, Ishan, et al.
Veröffentlicht: (2024)
von: Kavathekar, Ishan, et al.
Veröffentlicht: (2024)
Benchmarking VLMs' Reasoning About Persuasive Atypical Images
von: Malakouti, Sina, et al.
Veröffentlicht: (2024)
von: Malakouti, Sina, et al.
Veröffentlicht: (2024)
In-Context Learning through the Bayesian Prism
von: Panwar, Madhur, et al.
Veröffentlicht: (2023)
von: Panwar, Madhur, et al.
Veröffentlicht: (2023)
On the Quantization Robustness of Diffusion Language Models in Coding Benchmarks
von: Gupta, Aarav, et al.
Veröffentlicht: (2026)
von: Gupta, Aarav, et al.
Veröffentlicht: (2026)
IL-TUR: Benchmark for Indian Legal Text Understanding and Reasoning
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
von: Joshi, Abhinav, et al.
Veröffentlicht: (2024)
Thinking Fair and Slow: On the Efficacy of Structured Prompts for Debiasing Language Models
von: Furniturewala, Shaz, et al.
Veröffentlicht: (2024)
von: Furniturewala, Shaz, et al.
Veröffentlicht: (2024)
Provably Robust DPO: Aligning Language Models with Noisy Feedback
von: Chowdhury, Sayak Ray, et al.
Veröffentlicht: (2024)
von: Chowdhury, Sayak Ray, et al.
Veröffentlicht: (2024)
sign.mt: Real-Time Multilingual Sign Language Translation Application
von: Moryossef, Amit
Veröffentlicht: (2023)
von: Moryossef, Amit
Veröffentlicht: (2023)
CurLL: A Developmental Framework to Evaluate Continual Learning in Language Models
von: Kalyan, Pavan, et al.
Veröffentlicht: (2025)
von: Kalyan, Pavan, et al.
Veröffentlicht: (2025)
VOLTAGE: A Versatile Contrastive Learning based OCR Methodology for ultra low-resource scripts through Auto Glyph Feature Extraction
von: Sharma, Prawaal, et al.
Veröffentlicht: (2025)
von: Sharma, Prawaal, et al.
Veröffentlicht: (2025)
Application Specific Compression of Deep Learning Models
von: Rai, Rohit Raj, et al.
Veröffentlicht: (2024)
von: Rai, Rohit Raj, et al.
Veröffentlicht: (2024)
CoCoA: Confidence and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Models
von: Khandelwal, Anant, et al.
Veröffentlicht: (2025)
von: Khandelwal, Anant, et al.
Veröffentlicht: (2025)
HistoryBankQA: Multilingual Temporal Question Answering on Historical Events
von: Mandal, Biswadip, et al.
Veröffentlicht: (2025)
von: Mandal, Biswadip, et al.
Veröffentlicht: (2025)
Exploring Continual Fine-Tuning for Enhancing Language Ability in Large Language Model
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
von: Aggarwal, Divyanshu, et al.
Veröffentlicht: (2024)
Causal Order: The Key to Leveraging Imperfect Experts in Causal Inference
von: Vashishtha, Aniket, et al.
Veröffentlicht: (2023)
von: Vashishtha, Aniket, et al.
Veröffentlicht: (2023)
Cognitive Fatigue in Autoregressive Transformers: Formalization and Measurement
von: Marwah, Riju, et al.
Veröffentlicht: (2026)
von: Marwah, Riju, et al.
Veröffentlicht: (2026)
HORIZON: A Benchmark for In-the-wild User Behaviour Modeling
von: Goel, Arnav, et al.
Veröffentlicht: (2026)
von: Goel, Arnav, et al.
Veröffentlicht: (2026)
Multi-Preference Optimization: Generalizing DPO via Set-Level Contrasts
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
von: Gupta, Taneesh, et al.
Veröffentlicht: (2024)
AccessEval: Benchmarking Disability Bias in Large Language Models
von: Panda, Srikant, et al.
Veröffentlicht: (2025)
von: Panda, Srikant, et al.
Veröffentlicht: (2025)
InversionView: A General-Purpose Method for Reading Information from Neural Activations
von: Huang, Xinting, et al.
Veröffentlicht: (2024)
von: Huang, Xinting, et al.
Veröffentlicht: (2024)
Real-Time Multilingual Sign Language Processing
von: Moryossef, Amit
Veröffentlicht: (2024)
von: Moryossef, Amit
Veröffentlicht: (2024)
NIM: Neuro-symbolic Ideographic Metalanguage for Inclusive Communication
von: Sharma, Prawaal, et al.
Veröffentlicht: (2025)
von: Sharma, Prawaal, et al.
Veröffentlicht: (2025)
NeuroLit Navigator: A Neurosymbolic Approach to Scholarly Article Searches for Systematic Reviews
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2025)
von: Khandelwal, Vedant, et al.
Veröffentlicht: (2025)
DeduCE: Deductive Consistency as a Framework to Evaluate LLM Reasoning
von: Pandey, Atharva, et al.
Veröffentlicht: (2025)
von: Pandey, Atharva, et al.
Veröffentlicht: (2025)
A fully automated and scalable Parallel Data Augmentation for Low Resource Languages using Image and Text Analytics
von: Sharma, Prawaal, et al.
Veröffentlicht: (2025)
von: Sharma, Prawaal, et al.
Veröffentlicht: (2025)
Tricking LLMs into Disobedience: Formalizing, Analyzing, and Detecting Jailbreaks
von: Rao, Abhinav, et al.
Veröffentlicht: (2023)
von: Rao, Abhinav, et al.
Veröffentlicht: (2023)
Inferring Functionality of Attention Heads from their Parameters
von: Elhelo, Amit, et al.
Veröffentlicht: (2024)
von: Elhelo, Amit, et al.
Veröffentlicht: (2024)
SignBank+: Preparing a Multilingual Sign Language Dataset for Machine Translation Using Large Language Models
von: Moryossef, Amit, et al.
Veröffentlicht: (2023)
von: Moryossef, Amit, et al.
Veröffentlicht: (2023)
RE-IMAGINE: Symbolic Benchmark Synthesis for Reasoning Evaluation
von: Xu, Xinnuo, et al.
Veröffentlicht: (2025)
von: Xu, Xinnuo, et al.
Veröffentlicht: (2025)
Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks
von: Ganguly, Debargha, et al.
Veröffentlicht: (2025)
von: Ganguly, Debargha, et al.
Veröffentlicht: (2025)
Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
von: Gupta, Shashank, et al.
Veröffentlicht: (2023)
von: Gupta, Shashank, et al.
Veröffentlicht: (2023)
"Define Your Terms" : Enhancing Efficient Offensive Speech Classification with Definition
von: Nghiem, Huy, et al.
Veröffentlicht: (2024)
von: Nghiem, Huy, et al.
Veröffentlicht: (2024)
Teaching Transformers Causal Reasoning through Axiomatic Training
von: Vashishtha, Aniket, et al.
Veröffentlicht: (2024)
von: Vashishtha, Aniket, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
FrugalRAG: Less is More in RL Finetuning for Multi-Hop Question Answering
von: Java, Abhinav, et al.
Veröffentlicht: (2025) -
Plan*RAG: Efficient Test-Time Planning for Retrieval Augmented Generation
von: Verma, Prakhar, et al.
Veröffentlicht: (2024) -
Achieving Limited Adaptivity for Multinomial Logistic Bandits
von: Midigeshi, Sukruta Prakash, et al.
Veröffentlicht: (2025) -
NICE: To Optimize In-Context Examples or Not?
von: Srivastava, Pragya, et al.
Veröffentlicht: (2024) -
Task Facet Learning: A Structured Approach to Prompt Optimization
von: Juneja, Gurusha, et al.
Veröffentlicht: (2024)