The Geometric Anatomy of Capability Acquisition in Transformers
Fuente:
arXiv
Saved in:
| Main Author: | Billa, Jayadev |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
Predicting Where Steering Vectors Succeed
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
Algorithmic Capabilities of Random Transformers
by: Zhong, Ziqian, et al.
Published: (2024)
by: Zhong, Ziqian, et al.
Published: (2024)
MorphNAS: Differentiable Architecture Search for Morphologically-Aware Multilingual NER
by: Devadiga, Prathamesh, et al.
Published: (2025)
by: Devadiga, Prathamesh, et al.
Published: (2025)
AcquisitionSynthesis: Targeted Data Generation using Acquisition Functions
by: Agarwal, Ishika, et al.
Published: (2026)
by: Agarwal, Ishika, et al.
Published: (2026)
Schoenfeld's Anatomy of Mathematical Reasoning by Language Models
by: Li, Ming, et al.
Published: (2025)
by: Li, Ming, et al.
Published: (2025)
Anatomy of an Idiom: Tracing Non-Compositionality in Language Models
by: Gomes, Andrew
Published: (2025)
by: Gomes, Andrew
Published: (2025)
Decomposing the Depth Profile of Fine-Tuning
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
Data Mixing Can Induce Phase Transitions in Knowledge Acquisition
by: Gu, Xinran, et al.
Published: (2025)
by: Gu, Xinran, et al.
Published: (2025)
Human Inspired Progressive Alignment and Comparative Learning for Grounded Word Acquisition
by: Bao, Yuwei, et al.
Published: (2023)
by: Bao, Yuwei, et al.
Published: (2023)
When Audio-LLMs Don't Listen: A Cross-Linguistic Study of Modality Arbitration
by: Billa, Jayadev
Published: (2026)
by: Billa, Jayadev
Published: (2026)
Peering Through Preferences: Unraveling Feedback Acquisition for Aligning Large Language Models
by: Bansal, Hritik, et al.
Published: (2023)
by: Bansal, Hritik, et al.
Published: (2023)
Geometric-disentangelment Unlearning
by: Zhou, Duo, et al.
Published: (2025)
by: Zhou, Duo, et al.
Published: (2025)
Quantifying the Capabilities of LLMs across Scale and Precision
by: Badshah, Sher, et al.
Published: (2024)
by: Badshah, Sher, et al.
Published: (2024)
On Calibration of Large Language Models: From Response To Capability
by: Yang, Sin-Han, et al.
Published: (2026)
by: Yang, Sin-Han, et al.
Published: (2026)
Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
by: Zhang, Hanlin, et al.
Published: (2026)
by: Zhang, Hanlin, et al.
Published: (2026)
AI Scientists Fail Without Strong Implementation Capability
by: Zhu, Minjun, et al.
Published: (2025)
by: Zhu, Minjun, et al.
Published: (2025)
Exploring and Benchmarking the Planning Capabilities of Large Language Models
by: Bohnet, Bernd, et al.
Published: (2024)
by: Bohnet, Bernd, et al.
Published: (2024)
Bridging the Knowledge Void: Inference-time Acquisition of Unfamiliar Programming Languages for Coding Tasks
by: Shen, Chen, et al.
Published: (2026)
by: Shen, Chen, et al.
Published: (2026)
Unlocking Reasoning Capabilities in LLMs via Reinforcement Learning Exploration
by: Deng, Wenhao, et al.
Published: (2025)
by: Deng, Wenhao, et al.
Published: (2025)
ALPINE: Unveiling the Planning Capability of Autoregressive Learning in Language Models
by: Wang, Siwei, et al.
Published: (2024)
by: Wang, Siwei, et al.
Published: (2024)
Automated Capability Discovery via Foundation Model Self-Exploration
by: Lu, Cong, et al.
Published: (2025)
by: Lu, Cong, et al.
Published: (2025)
ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
by: Karger, Ezra, et al.
Published: (2024)
by: Karger, Ezra, et al.
Published: (2024)
ModelGPT: Unleashing LLM's Capabilities for Tailored Model Generation
by: Tang, Zihao, et al.
Published: (2024)
by: Tang, Zihao, et al.
Published: (2024)
What is it for a Machine Learning Model to Have a Capability?
by: Harding, Jacqueline, et al.
Published: (2024)
by: Harding, Jacqueline, et al.
Published: (2024)
How Numerical Precision Affects Arithmetical Reasoning Capabilities of LLMs
by: Feng, Guhao, et al.
Published: (2024)
by: Feng, Guhao, et al.
Published: (2024)
Enhancing Reasoning Capabilities in SLMs with Reward Guided Dataset Distillation
by: Padarha, Shreyansh
Published: (2025)
by: Padarha, Shreyansh
Published: (2025)
Assessing the Impact of Prompting Methods on ChatGPT's Mathematical Capabilities
by: Chen, Yuhao, et al.
Published: (2023)
by: Chen, Yuhao, et al.
Published: (2023)
Bringing Up a Bilingual BabyLM: Investigating Multilingual Language Acquisition Using Small-Scale Models
by: Zeng, Linda, et al.
Published: (2026)
by: Zeng, Linda, et al.
Published: (2026)
Breaking the Capability Ceiling of LLM Post-Training by Reintroducing Markov States
by: Yuan, Yurun, et al.
Published: (2026)
by: Yuan, Yurun, et al.
Published: (2026)
EffGen: Enabling Small Language Models as Capable Autonomous Agents
by: Srivastava, Gaurav, et al.
Published: (2026)
by: Srivastava, Gaurav, et al.
Published: (2026)
Language Diffusion Models are Associative Memories Capable of Retrieving Unseen Data
by: Pham, Bao, et al.
Published: (2026)
by: Pham, Bao, et al.
Published: (2026)
Counterfactual Evaluation Reveals Hidden Capability Profiles in Clinical LLMs and Agents
by: Turk, Matt
Published: (2026)
by: Turk, Matt
Published: (2026)
LimiX: Unleashing Structured-Data Modeling Capability for Generalist Intelligence
by: Zhang, Xingxuan, et al.
Published: (2025)
by: Zhang, Xingxuan, et al.
Published: (2025)
CityBench: Evaluating the Capabilities of Large Language Models for Urban Tasks
by: Feng, Jie, et al.
Published: (2024)
by: Feng, Jie, et al.
Published: (2024)
Disentangling Logic: The Role of Context in Large Language Model Reasoning Capabilities
by: Hua, Wenyue, et al.
Published: (2024)
by: Hua, Wenyue, et al.
Published: (2024)
Capability Instruction Tuning: A New Paradigm for Dynamic LLM Routing
by: Zhang, Yi-Kai, et al.
Published: (2025)
by: Zhang, Yi-Kai, et al.
Published: (2025)
STEM: Efficient Relative Capability Evaluation of LLMs through Structured Transition Samples
by: Hu, Haiquan, et al.
Published: (2025)
by: Hu, Haiquan, et al.
Published: (2025)
Scalpel vs. Hammer: GRPO Amplifies Existing Capabilities, SFT Replaces Them
by: Rajani, Neel, et al.
Published: (2025)
by: Rajani, Neel, et al.
Published: (2025)
Similar Items
-
Modality Collapse as Mismatched Decoding: Information-Theoretic Limits of Multimodal LLMs
by: Billa, Jayadev
Published: (2026) -
Predicting Where Steering Vectors Succeed
by: Billa, Jayadev
Published: (2026) -
The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR$\rightarrow$LLM Pipelines?
by: Billa, Jayadev
Published: (2026) -
Algorithmic Capabilities of Random Transformers
by: Zhong, Ziqian, et al.
Published: (2024) -
MorphNAS: Differentiable Architecture Search for Morphologically-Aware Multilingual NER
by: Devadiga, Prathamesh, et al.
Published: (2025)