Law of the Weakest Link: Cross Capabilities of Large Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Zhong, Ming, Zhang, Aston, Wang, Xuewei, Hou, Rui, Xiong, Wenhan, Zhu, Chenguang, Chen, Zhengxing, Tan, Liang, Bi, Chloe, Lewis, Mike, Popuri, Sravya, Narang, Sharan, Kambadur, Melanie, Mahajan, Dhruv, Edunov, Sergey, Han, Jiawei, van der Maaten, Laurens |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Self-Generated Critiques Boost Reward Modeling for Language Models
by: Yu, Yue, et al.
Published: (2024)
by: Yu, Yue, et al.
Published: (2024)
A Systematic Examination of Preference Learning through the Lens of Instruction-Following
by: Kim, Joongwon, et al.
Published: (2024)
by: Kim, Joongwon, et al.
Published: (2024)
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
by: Roberts, Nicholas, et al.
Published: (2025)
by: Roberts, Nicholas, et al.
Published: (2025)
Parallel-SFT: Improving Zero-Shot Cross-Programming-Language Transfer for Code RL
by: Wu, Zhaofeng, et al.
Published: (2026)
by: Wu, Zhaofeng, et al.
Published: (2026)
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis
by: Peng, Yifan, et al.
Published: (2024)
by: Peng, Yifan, et al.
Published: (2024)
MSLM-S2ST: A Multitask Speech Language Model for Textless Speech-to-Speech Translation with Speaker Style Preservation
by: Peng, Yifan, et al.
Published: (2024)
by: Peng, Yifan, et al.
Published: (2024)
Investigating Decoder-only Large Language Models for Speech-to-text Translation
by: Huang, Chao-Wei, et al.
Published: (2024)
by: Huang, Chao-Wei, et al.
Published: (2024)
Guarantees of confidentiality via Hammersley-Chapman-Robbins bounds
by: Chaudhuri, Kamalika, et al.
Published: (2024)
by: Chaudhuri, Kamalika, et al.
Published: (2024)
Infinite Problem Generator: Verifiably Scaling Physics Reasoning Data with Agentic Workflows
by: Sharan, Aditya, et al.
Published: (2026)
by: Sharan, Aditya, et al.
Published: (2026)
Investigating Spatial Attention Bias in Vision-Language Models
by: Chaudhary, Aryan, et al.
Published: (2025)
by: Chaudhary, Aryan, et al.
Published: (2025)
Evolution analysis of software quality metrics in an open-source java project: A case study on TestNG
by: Sambaturu, Venkata Sai Sravya
Published: (2025)
by: Sambaturu, Venkata Sai Sravya
Published: (2025)
A Weakest Precondition Calculus for Programs and Linear Temporal Specifications
by: Ernst, Gidon
Published: (2026)
by: Ernst, Gidon
Published: (2026)
Correlating and Predicting Human Evaluations of Language Models from Natural Language Processing Benchmarks
by: Schaeffer, Rylan, et al.
Published: (2025)
by: Schaeffer, Rylan, et al.
Published: (2025)
Secure multiparty computations in floating-point arithmetic
by: Guo, Chuan, et al.
Published: (2020)
by: Guo, Chuan, et al.
Published: (2020)
Higher-Order Weakest Precondition Transformers via a CPS Transformation
by: Kura, Satoshi
Published: (2023)
by: Kura, Satoshi
Published: (2023)
Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following
by: He, Yun, et al.
Published: (2024)
by: He, Yun, et al.
Published: (2024)
You Only Look at Screens: Multimodal Chain-of-Action Agents
by: Zhang, Zhuosheng, et al.
Published: (2023)
by: Zhang, Zhuosheng, et al.
Published: (2023)
Dual Forgetting Operators in the Context of Weakest Sufficient and Strongest Necessary Conditions
by: Doherty, Patrick, et al.
Published: (2023)
by: Doherty, Patrick, et al.
Published: (2023)
Weakest Bidder Types and New Core-Selecting Combinatorial Auctions
by: Prasad, Siddharth, et al.
Published: (2025)
by: Prasad, Siddharth, et al.
Published: (2025)
A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains
by: Jacovi, Alon, et al.
Published: (2024)
by: Jacovi, Alon, et al.
Published: (2024)
The Weakest Link: Library Catalogs.
by: Young, Terrence E., Jr.
Published: (2002)
by: Young, Terrence E., Jr.
Published: (2002)
GeoDE: a Geographically Diverse Evaluation Dataset for Object Recognition
by: Ramaswamy, Vikram V., et al.
Published: (2023)
by: Ramaswamy, Vikram V., et al.
Published: (2023)
A Hybrid Supervised-LLM Pipeline for Actionable Suggestion Mining in Unstructured Customer Reviews
by: Trivedi, Aakash, et al.
Published: (2026)
by: Trivedi, Aakash, et al.
Published: (2026)
Reinforcing the Weakest Links: Modernizing SIENA with Targeted Deep Learning Integration
by: Raciti, Riccardo, et al.
Published: (2026)
by: Raciti, Riccardo, et al.
Published: (2026)
Balancing Enhancement, Harmlessness, and General Capabilities: Enhancing Conversational LLMs with Direct RLHF
by: Zheng, Chen, et al.
Published: (2024)
by: Zheng, Chen, et al.
Published: (2024)
LLMs and Fuzzing in Tandem: A New Approach to Automatically Generating Weakest Preconditions
by: King, Daragh, et al.
Published: (2025)
by: King, Daragh, et al.
Published: (2025)
A Neurosymbolic Approach to Loop Invariant Generation via Weakest Precondition Reasoning
by: King, Daragh, et al.
Published: (2025)
by: King, Daragh, et al.
Published: (2025)
Health Insurance Coverage Rule Interpretation Corpus: Law, Policy, and Medical Guidance for Health Insurance Coverage Understanding
by: Gartner, Mike
Published: (2025)
by: Gartner, Mike
Published: (2025)
From Facts to Conclusions : Integrating Deductive Reasoning in Retrieval-Augmented LLMs
by: Mishra, Shubham, et al.
Published: (2025)
by: Mishra, Shubham, et al.
Published: (2025)
Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law
by: Ge, Qiming, et al.
Published: (2025)
by: Ge, Qiming, et al.
Published: (2025)
What is Ethical: AIHED Driving Humans or Human-Driven AIHED? A Conceptual Framework enabling the Ethos of AI-driven Higher education
by: Mahajan, Prashant
Published: (2025)
by: Mahajan, Prashant
Published: (2025)
Using Counterfactual Tasks to Evaluate the Generality of Analogical Reasoning in Large Language Models
by: Lewis, Martha, et al.
Published: (2024)
by: Lewis, Martha, et al.
Published: (2024)
Beyond Sentiment: A Multi-Agent Pipeline for Actionable Business Advice from Reviews
by: Bhandari, Kartikey Singh, et al.
Published: (2026)
by: Bhandari, Kartikey Singh, et al.
Published: (2026)
IndicDB -- Benchmarking Multilingual Text-to-SQL Capabilities in Indian Languages
by: Dawar, Aviral, et al.
Published: (2026)
by: Dawar, Aviral, et al.
Published: (2026)
Multimodal ML: Quantifying the Improvement of Calorie Estimation Through Image-Text Pairs
by: Narang, Arya
Published: (2025)
by: Narang, Arya
Published: (2025)
Lights, Camera, Consistency: A Multistage Pipeline for Character-Stable AI Video Stories
by: Jain, Chayan, et al.
Published: (2025)
by: Jain, Chayan, et al.
Published: (2025)
Domain-Partitioned Hybrid RAG for Legal Reasoning: Toward Modular and Explainable Legal AI for India
by: Goel, Rakshita, et al.
Published: (2025)
by: Goel, Rakshita, et al.
Published: (2025)
NOMAD Projection
by: Duderstadt, Brandon, et al.
Published: (2025)
by: Duderstadt, Brandon, et al.
Published: (2025)
Ulterior Motives: Detecting Misaligned Reasoning in Continuous Thought Models
by: Ramjee, Sharan
Published: (2026)
by: Ramjee, Sharan
Published: (2026)
AI Family Integration Index (AFII): Benchmarking a New Global Readiness for AI as Family
by: Mahajan, Prashant
Published: (2025)
by: Mahajan, Prashant
Published: (2025)
Similar Items
-
Self-Generated Critiques Boost Reward Modeling for Language Models
by: Yu, Yue, et al.
Published: (2024) -
A Systematic Examination of Preference Learning through the Lens of Instruction-Following
by: Kim, Joongwon, et al.
Published: (2024) -
Compute Optimal Scaling of Skills: Knowledge vs Reasoning
by: Roberts, Nicholas, et al.
Published: (2025) -
Parallel-SFT: Improving Zero-Shot Cross-Programming-Language Transfer for Code RL
by: Wu, Zhaofeng, et al.
Published: (2026) -
An Empirical Study of Speech Language Models for Prompt-Conditioned Speech Synthesis
by: Peng, Yifan, et al.
Published: (2024)