Function Induction and Task Generalization: An Interpretability Study with Off-by-One Addition
Fuente:
arXiv
Salvato in:
| Autori principali: | Ye, Qinyuan, Jia, Robin, Ren, Xiang |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Stress-Testing Long-Context Language Models with Lifelong ICL and Task Haystack
di: Xu, Xiaoyue, et al.
Pubblicazione: (2024)
di: Xu, Xiaoyue, et al.
Pubblicazione: (2024)
Interpretable Next-token Prediction via the Generalized Induction Head
di: Kim, Eunji, et al.
Pubblicazione: (2024)
di: Kim, Eunji, et al.
Pubblicazione: (2024)
GPTQT: Quantize Large Language Models Twice to Push the Efficiency
di: Guo, Yipin, et al.
Pubblicazione: (2024)
di: Guo, Yipin, et al.
Pubblicazione: (2024)
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
di: Yan, Tianyi Lorena, et al.
Pubblicazione: (2025)
di: Yan, Tianyi Lorena, et al.
Pubblicazione: (2025)
Prompt Engineering a Prompt Engineer
di: Ye, Qinyuan, et al.
Pubblicazione: (2023)
di: Ye, Qinyuan, et al.
Pubblicazione: (2023)
Demystifying the Accuracy-Interpretability Trade-Off: A Case Study of Inferring Ratings from Reviews
di: Atrey, Pranjal, et al.
Pubblicazione: (2025)
di: Atrey, Pranjal, et al.
Pubblicazione: (2025)
Interpretable Steering of Large Language Models with Feature Guided Activation Additions
di: Soo, Samuel, et al.
Pubblicazione: (2025)
di: Soo, Samuel, et al.
Pubblicazione: (2025)
Assessing Episodic Memory in LLMs with Sequence Order Recall Tasks
di: Pink, Mathis, et al.
Pubblicazione: (2024)
di: Pink, Mathis, et al.
Pubblicazione: (2024)
Mechanistic Interpretability of GPT-like Models on Summarization Tasks
di: Mishra, Anurag
Pubblicazione: (2025)
di: Mishra, Anurag
Pubblicazione: (2025)
FECT: Factuality Evaluation of Interpretive AI-Generated Claims in Contact Center Conversation Transcripts
di: Shin, Hagyeong, et al.
Pubblicazione: (2025)
di: Shin, Hagyeong, et al.
Pubblicazione: (2025)
Task-Circuit Quantization: Leveraging Knowledge Localization and Interpretability for Compression
di: Xiao, Hanqi, et al.
Pubblicazione: (2025)
di: Xiao, Hanqi, et al.
Pubblicazione: (2025)
DETAIL: Task DEmonsTration Attribution for Interpretable In-context Learning
di: Zhou, Zijian, et al.
Pubblicazione: (2024)
di: Zhou, Zijian, et al.
Pubblicazione: (2024)
Combining Induction and Transduction for Abstract Reasoning
di: Li, Wen-Ding, et al.
Pubblicazione: (2024)
di: Li, Wen-Ding, et al.
Pubblicazione: (2024)
Universal Response and Emergence of Induction in LLMs
di: Luick, Niclas
Pubblicazione: (2024)
di: Luick, Niclas
Pubblicazione: (2024)
Team QUST at SemEval-2023 Task 3: A Comprehensive Study of Monolingual and Multilingual Approaches for Detecting Online News Genre, Framing and Persuasion Techniques
di: Jiang, Ye
Pubblicazione: (2023)
di: Jiang, Ye
Pubblicazione: (2023)
Induction Signatures Are Not Enough: A Matched-Compute Study of Load-Bearing Structure in In-Context Learning
di: Sabry, Mohammed, et al.
Pubblicazione: (2025)
di: Sabry, Mohammed, et al.
Pubblicazione: (2025)
Revisiting the Test-Time Scaling of o1-like Models: Do they Truly Possess Test-Time Scaling Capabilities?
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2025)
di: Zeng, Zhiyuan, et al.
Pubblicazione: (2025)
Large Language Models as Computable Approximations to Solomonoff Induction
di: Wan, Jun, et al.
Pubblicazione: (2025)
di: Wan, Jun, et al.
Pubblicazione: (2025)
Towards Efficient Neurally-Guided Program Induction for ARC-AGI
di: Ouellette, Simon
Pubblicazione: (2024)
di: Ouellette, Simon
Pubblicazione: (2024)
Re-evaluating the Need for Multimodal Signals in Unsupervised Grammar Induction
di: Li, Boyi, et al.
Pubblicazione: (2022)
di: Li, Boyi, et al.
Pubblicazione: (2022)
Language Models Use Trigonometry to Do Addition
di: Kantamneni, Subhash, et al.
Pubblicazione: (2025)
di: Kantamneni, Subhash, et al.
Pubblicazione: (2025)
Stream Aligner: Efficient Sentence-Level Alignment via Distribution Induction
di: Lou, Hantao, et al.
Pubblicazione: (2025)
di: Lou, Hantao, et al.
Pubblicazione: (2025)
Plan Optimization to Bilingual Dictionary Induction for Low-Resource Language Families
di: Nasution, Arbi Haza, et al.
Pubblicazione: (2020)
di: Nasution, Arbi Haza, et al.
Pubblicazione: (2020)
SepLLM: Accelerate Large Language Models by Compressing One Segment into One Separator
di: Chen, Guoxuan, et al.
Pubblicazione: (2024)
di: Chen, Guoxuan, et al.
Pubblicazione: (2024)
Understanding In-context Learning of Addition via Activation Subspaces
di: Hu, Xinyan, et al.
Pubblicazione: (2025)
di: Hu, Xinyan, et al.
Pubblicazione: (2025)
Steering Llama 2 via Contrastive Activation Addition
di: Panickssery, Nina, et al.
Pubblicazione: (2023)
di: Panickssery, Nina, et al.
Pubblicazione: (2023)
Interpreting and Controlling LLM Reasoning through Integrated Policy Gradient
di: Li, Changming, et al.
Pubblicazione: (2026)
di: Li, Changming, et al.
Pubblicazione: (2026)
Granite-Function Calling Model: Introducing Function Calling Abilities via Multi-task Learning of Granular Tasks
di: Abdelaziz, Ibrahim, et al.
Pubblicazione: (2024)
di: Abdelaziz, Ibrahim, et al.
Pubblicazione: (2024)
Instruction Diversity Drives Generalization To Unseen Tasks
di: Zhang, Dylan, et al.
Pubblicazione: (2024)
di: Zhang, Dylan, et al.
Pubblicazione: (2024)
AutoLibra: Agent Metric Induction from Open-Ended Human Feedback
di: Zhu, Hao, et al.
Pubblicazione: (2025)
di: Zhu, Hao, et al.
Pubblicazione: (2025)
TracrBench: Generating Interpretability Testbeds with Large Language Models
di: Thurnherr, Hannes, et al.
Pubblicazione: (2024)
di: Thurnherr, Hannes, et al.
Pubblicazione: (2024)
A General Framework for Producing Interpretable Semantic Text Embeddings
di: Sun, Yiqun, et al.
Pubblicazione: (2024)
di: Sun, Yiqun, et al.
Pubblicazione: (2024)
Learning to Reason under Off-Policy Guidance
di: Yan, Jianhao, et al.
Pubblicazione: (2025)
di: Yan, Jianhao, et al.
Pubblicazione: (2025)
Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning
di: Zhao, Shiwan, et al.
Pubblicazione: (2026)
di: Zhao, Shiwan, et al.
Pubblicazione: (2026)
CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis
di: Chen, Junying, et al.
Pubblicazione: (2024)
di: Chen, Junying, et al.
Pubblicazione: (2024)
Transformers Learn to Achieve Second-Order Convergence Rates for In-Context Linear Regression
di: Fu, Deqing, et al.
Pubblicazione: (2023)
di: Fu, Deqing, et al.
Pubblicazione: (2023)
Interpretable Predictability-Based AI Text Detection: A Replication Study
di: Skurla, Adam, et al.
Pubblicazione: (2026)
di: Skurla, Adam, et al.
Pubblicazione: (2026)
Pretrained Generative Language Models as General Learning Frameworks for Sequence-Based Tasks
di: Fauber, Ben
Pubblicazione: (2024)
di: Fauber, Ben
Pubblicazione: (2024)
Exploring the Hidden Capacity of LLMs for One-Step Text Generation
di: Mezentsev, Gleb, et al.
Pubblicazione: (2025)
di: Mezentsev, Gleb, et al.
Pubblicazione: (2025)
A Dataset for Evaluating LLM-based Evaluation Functions for Research Question Extraction Task
di: Fujisaki, Yuya, et al.
Pubblicazione: (2024)
di: Fujisaki, Yuya, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Stress-Testing Long-Context Language Models with Lifelong ICL and Task Haystack
di: Xu, Xiaoyue, et al.
Pubblicazione: (2024) -
Interpretable Next-token Prediction via the Generalized Induction Head
di: Kim, Eunji, et al.
Pubblicazione: (2024) -
GPTQT: Quantize Large Language Models Twice to Push the Efficiency
di: Guo, Yipin, et al.
Pubblicazione: (2024) -
Promote, Suppress, Iterate: How Language Models Answer One-to-Many Factual Queries
di: Yan, Tianyi Lorena, et al.
Pubblicazione: (2025) -
Prompt Engineering a Prompt Engineer
di: Ye, Qinyuan, et al.
Pubblicazione: (2023)