Salvato in:
| Autori principali: | Jacovi, Alon, Bitton, Yonatan, Bohnet, Bernd, Herzig, Jonathan, Honovich, Or, Tseng, Michael, Collins, Michael, Aharoni, Roee, Geva, Mor |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2402.00559 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
di: Yona, Gal, et al.
Pubblicazione: (2024)
di: Yona, Gal, et al.
Pubblicazione: (2024)
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
di: Yona, Gal, et al.
Pubblicazione: (2024)
di: Yona, Gal, et al.
Pubblicazione: (2024)
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
di: Gekhman, Zorik, et al.
Pubblicazione: (2026)
di: Gekhman, Zorik, et al.
Pubblicazione: (2026)
Keep Guessing? When Considering Inference Scaling, Mind the Baselines
di: Yona, Gal, et al.
Pubblicazione: (2024)
di: Yona, Gal, et al.
Pubblicazione: (2024)
DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
di: Cattan, Arie, et al.
Pubblicazione: (2025)
di: Cattan, Arie, et al.
Pubblicazione: (2025)
CoverBench: A Challenging Benchmark for Complex Claim Verification
di: Jacovi, Alon, et al.
Pubblicazione: (2024)
di: Jacovi, Alon, et al.
Pubblicazione: (2024)
Marketing the Librarian: The Weakest Link in the Chain.
di: Kies, Cosette
Pubblicazione: (1989)
di: Kies, Cosette
Pubblicazione: (1989)
Weakest Link in the Chain: Security Vulnerabilities in Advanced Reasoning Models
di: Krishna, Arjun, et al.
Pubblicazione: (2025)
di: Krishna, Arjun, et al.
Pubblicazione: (2025)
FinChain: A Symbolic Benchmark for Verifiable Chain-of-Thought Financial Reasoning
di: Xie, Zhuohan, et al.
Pubblicazione: (2025)
di: Xie, Zhuohan, et al.
Pubblicazione: (2025)
DoubleDipper: Improving Long-Context LLMs via Context Recycling
di: Cattan, Arie, et al.
Pubblicazione: (2024)
di: Cattan, Arie, et al.
Pubblicazione: (2024)
Verifying Chain-of-Thought Reasoning via Its Computational Graph
di: Zhao, Zheng, et al.
Pubblicazione: (2025)
di: Zhao, Zheng, et al.
Pubblicazione: (2025)
On Learning Verifiers and Implications to Chain-of-Thought Reasoning
di: Balcan, Maria-Florina, et al.
Pubblicazione: (2025)
di: Balcan, Maria-Florina, et al.
Pubblicazione: (2025)
Measuring Chain of Thought Faithfulness by Unlearning Reasoning Steps
di: Tutek, Martin, et al.
Pubblicazione: (2025)
di: Tutek, Martin, et al.
Pubblicazione: (2025)
Evaluating Chain-of-Thought Reasoning through Reusability and Verifiability
di: Aggarwal, Shashank, et al.
Pubblicazione: (2026)
di: Aggarwal, Shashank, et al.
Pubblicazione: (2026)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
di: Katz, Shahar, et al.
Pubblicazione: (2024)
di: Katz, Shahar, et al.
Pubblicazione: (2024)
Accelerating the Global Aggregation of Local Explanations
di: Mor, Alon, et al.
Pubblicazione: (2023)
di: Mor, Alon, et al.
Pubblicazione: (2023)
MM-Verify: Enhancing Multimodal Reasoning with Chain-of-Thought Verification
di: Sun, Linzhuang, et al.
Pubblicazione: (2025)
di: Sun, Linzhuang, et al.
Pubblicazione: (2025)
NL-Eye: Abductive NLI for Images
di: Ventura, Mor, et al.
Pubblicazione: (2024)
di: Ventura, Mor, et al.
Pubblicazione: (2024)
Universal Jailbreak Suffixes Are Strong Attention Hijackers
di: Ben-Tov, Matan, et al.
Pubblicazione: (2025)
di: Ben-Tov, Matan, et al.
Pubblicazione: (2025)
Multilingual Instruction Tuning With Just a Pinch of Multilinguality
di: Shaham, Uri, et al.
Pubblicazione: (2024)
di: Shaham, Uri, et al.
Pubblicazione: (2024)
mFACE: Multilingual Summarization with Factual Consistency Evaluation
di: Aharoni, Roee, et al.
Pubblicazione: (2022)
di: Aharoni, Roee, et al.
Pubblicazione: (2022)
Representation Surgery: Theory and Practice of Affine Steering
di: Singh, Shashwat, et al.
Pubblicazione: (2024)
di: Singh, Shashwat, et al.
Pubblicazione: (2024)
EgoCoT-Bench: Benchmarking Grounded and Verifiable Operation-Centric Chain of Thought Reasoning for MLLMs
di: Dai, Yang, et al.
Pubblicazione: (2026)
di: Dai, Yang, et al.
Pubblicazione: (2026)
Typed Chain-of-Thought: A Curry-Howard Framework for Verifying LLM Reasoning
di: Perrier, Elija
Pubblicazione: (2025)
di: Perrier, Elija
Pubblicazione: (2025)
Latent Reasoning with Supervised Thinking States
di: Amos, Ido, et al.
Pubblicazione: (2026)
di: Amos, Ido, et al.
Pubblicazione: (2026)
Editor's Choice: Evaluating Abstract Intent in Image Editing through Atomic Entity Analysis
di: Ventura, Mor, et al.
Pubblicazione: (2026)
di: Ventura, Mor, et al.
Pubblicazione: (2026)
TACT: Advancing Complex Aggregative Reasoning with Information Extraction Tools
di: Caciularu, Avi, et al.
Pubblicazione: (2024)
di: Caciularu, Avi, et al.
Pubblicazione: (2024)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
di: Mitra, Chancharik, et al.
Pubblicazione: (2023)
di: Mitra, Chancharik, et al.
Pubblicazione: (2023)
How Well Can Reasoning Models Identify and Recover from Unhelpful Thoughts?
di: Yang, Sohee, et al.
Pubblicazione: (2025)
di: Yang, Sohee, et al.
Pubblicazione: (2025)
EditInspector: A Benchmark for Evaluation of Text-Guided Image Edits
di: Yosef, Ron, et al.
Pubblicazione: (2025)
di: Yosef, Ron, et al.
Pubblicazione: (2025)
Honest Students from Untrusted Teachers: Learning an Interpretable Question-Answering Pipeline from a Pretrained Language Model
di: Eisenstein, Jacob, et al.
Pubblicazione: (2022)
di: Eisenstein, Jacob, et al.
Pubblicazione: (2022)
CoRGI: Verified Chain-of-Thought Reasoning with Post-hoc Visual Grounding
di: Yi, Shixin, et al.
Pubblicazione: (2025)
di: Yi, Shixin, et al.
Pubblicazione: (2025)
Does Fine-Tuning LLMs on New Knowledge Encourage Hallucinations?
di: Gekhman, Zorik, et al.
Pubblicazione: (2024)
di: Gekhman, Zorik, et al.
Pubblicazione: (2024)
Estimating Knowledge in Large Language Models Without Generating a Single Token
di: Gottesman, Daniela, et al.
Pubblicazione: (2024)
di: Gottesman, Daniela, et al.
Pubblicazione: (2024)
Inferring Functionality of Attention Heads from their Parameters
di: Elhelo, Amit, et al.
Pubblicazione: (2024)
di: Elhelo, Amit, et al.
Pubblicazione: (2024)
The Weakest Link: Library Catalogs.
di: Young, Terrence E., Jr.
Pubblicazione: (2002)
di: Young, Terrence E., Jr.
Pubblicazione: (2002)
Generating Verifiable Chain of Thoughts from Exection-Traces
di: Thakur, Shailja, et al.
Pubblicazione: (2025)
di: Thakur, Shailja, et al.
Pubblicazione: (2025)
Fractured Chain-of-Thought Reasoning
di: Liao, Baohao, et al.
Pubblicazione: (2025)
di: Liao, Baohao, et al.
Pubblicazione: (2025)
GeoChain: Multimodal Chain-of-Thought for Geographic Reasoning
di: Yerramilli, Sahiti, et al.
Pubblicazione: (2025)
di: Yerramilli, Sahiti, et al.
Pubblicazione: (2025)
LongCoT: Benchmarking Long-Horizon Chain-of-Thought Reasoning
di: Motwani, Sumeet Ramesh, et al.
Pubblicazione: (2026)
di: Motwani, Sumeet Ramesh, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Narrowing the Knowledge Evaluation Gap: Open-Domain Question Answering with Multi-Granularity Answers
di: Yona, Gal, et al.
Pubblicazione: (2024) -
Can Large Language Models Faithfully Express Their Intrinsic Uncertainty in Words?
di: Yona, Gal, et al.
Pubblicazione: (2024) -
Thinking to Recall: How Reasoning Unlocks Parametric Knowledge in LLMs
di: Gekhman, Zorik, et al.
Pubblicazione: (2026) -
Keep Guessing? When Considering Inference Scaling, Mind the Baselines
di: Yona, Gal, et al.
Pubblicazione: (2024) -
DRAGged into Conflicts: Detecting and Addressing Conflicting Sources in Search-Augmented LLMs
di: Cattan, Arie, et al.
Pubblicazione: (2025)