Oolong: Investigating What Makes Transfer Learning Hard with Controlled Studies
Fuente:
arXiv
Saved in:
| Main Authors: | Wu, Zhengxuan, Tamkin, Alex, Papadimitriou, Isabel |
|---|---|
| Format: | Preprint |
| Published: |
2022
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
by: Bertsch, Amanda, et al.
Published: (2025)
by: Bertsch, Amanda, et al.
Published: (2025)
Vocabulary embeddings organize linguistic structure early in language model training
by: Papadimitriou, Isabel, et al.
Published: (2025)
by: Papadimitriou, Isabel, et al.
Published: (2025)
Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles
by: Bhattacharya, Antara Raaghavi, et al.
Published: (2025)
by: Bhattacharya, Antara Raaghavi, et al.
Published: (2025)
ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation
by: Wu, Zhengxuan, et al.
Published: (2023)
by: Wu, Zhengxuan, et al.
Published: (2023)
What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable Insights
by: Wen, Xin, et al.
Published: (2024)
by: Wen, Xin, et al.
Published: (2024)
Guiding Through Complexity: What Makes Good Supervision for Hard Math Reasoning Tasks?
by: He, Xuan, et al.
Published: (2024)
by: He, Xuan, et al.
Published: (2024)
Bayesian Preference Elicitation with Language Models
by: Handa, Kunal, et al.
Published: (2024)
by: Handa, Kunal, et al.
Published: (2024)
Simulated Language Acquisition in a Biologically Realistic Model of the Brain
by: Mitropolsky, Daniel, et al.
Published: (2025)
by: Mitropolsky, Daniel, et al.
Published: (2025)
AI "News" Content Farms Are Easy to Make and Hard to Detect: A Case Study in Italian
by: Puccetti, Giovanni, et al.
Published: (2024)
by: Puccetti, Giovanni, et al.
Published: (2024)
Multilinguality Does not Make Sense: Investigating Factors Behind Zero-Shot Transfer in Sense-Aware Tasks
by: Goworek, Roksana, et al.
Published: (2025)
by: Goworek, Roksana, et al.
Published: (2025)
LLMs Encode Harmfulness and Refusal Separately
by: Zhao, Jiachen, et al.
Published: (2025)
by: Zhao, Jiachen, et al.
Published: (2025)
Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
by: Denison, Carson, et al.
Published: (2024)
by: Denison, Carson, et al.
Published: (2024)
How AI Impacts Skill Formation
by: Shen, Judy Hanwen, et al.
Published: (2026)
by: Shen, Judy Hanwen, et al.
Published: (2026)
What Makes Good Instruction-Tuning Data? An In-Context Learning Perspective
by: Han, Guangzeng, et al.
Published: (2026)
by: Han, Guangzeng, et al.
Published: (2026)
ADAG: Automatically Describing Attribution Graphs
by: Arora, Aryaman, et al.
Published: (2026)
by: Arora, Aryaman, et al.
Published: (2026)
Interpretability at Scale: Identifying Causal Mechanisms in Alpaca
by: Wu, Zhengxuan, et al.
Published: (2023)
by: Wu, Zhengxuan, et al.
Published: (2023)
MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions
by: Zhong, Zexuan, et al.
Published: (2023)
by: Zhong, Zexuan, et al.
Published: (2023)
Improved Representation Steering for Language Models
by: Wu, Zhengxuan, et al.
Published: (2025)
by: Wu, Zhengxuan, et al.
Published: (2025)
Interpreting the linear structure of vision-language model embedding spaces
by: Papadimitriou, Isabel, et al.
Published: (2025)
by: Papadimitriou, Isabel, et al.
Published: (2025)
Bridging Cognition and Emotion: Empathy-Driven Multimodal Misinformation Detection
by: Wang, Zihan, et al.
Published: (2025)
by: Wang, Zihan, et al.
Published: (2025)
What Makes a Good Doctor Response? A Study on Text-Based Telemedicine
by: Cosma, Adrian, et al.
Published: (2026)
by: Cosma, Adrian, et al.
Published: (2026)
DataPuzzle: Breaking Free from the Hallucinated Promise of LLMs in Data Analysis
by: Zhang, Zhengxuan, et al.
Published: (2025)
by: Zhang, Zhengxuan, et al.
Published: (2025)
What Makes Language Models Good-enough?
by: Asami, Daiki, et al.
Published: (2024)
by: Asami, Daiki, et al.
Published: (2024)
What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
by: Jain, Samyak, et al.
Published: (2024)
by: Jain, Samyak, et al.
Published: (2024)
RAVEL: Evaluating Interpretability Methods on Disentangling Language Model Representations
by: Huang, Jing, et al.
Published: (2024)
by: Huang, Jing, et al.
Published: (2024)
Sakura at BEA 2026 Shared Task 1: What Makes Vocabulary Difficult?
by: Nohejl, Adam, et al.
Published: (2026)
by: Nohejl, Adam, et al.
Published: (2026)
Mission: Impossible Language Models
by: Kallini, Julie, et al.
Published: (2024)
by: Kallini, Julie, et al.
Published: (2024)
Language Model Circuits Are Sparse in the Neuron Basis
by: Arora, Aryaman, et al.
Published: (2026)
by: Arora, Aryaman, et al.
Published: (2026)
Sample Design Engineering: An Empirical Study of What Makes Good Downstream Fine-Tuning Samples for LLMs
by: Guo, Biyang, et al.
Published: (2024)
by: Guo, Biyang, et al.
Published: (2024)
What Makes a Good Natural Language Prompt?
by: Long, Do Xuan, et al.
Published: (2025)
by: Long, Do Xuan, et al.
Published: (2025)
What Makes Two Language Models Think Alike?
by: Salle, Jeanne, et al.
Published: (2024)
by: Salle, Jeanne, et al.
Published: (2024)
What Makes Math Word Problems Challenging for LLMs?
by: Srivatsa, KV Aditya, et al.
Published: (2024)
by: Srivatsa, KV Aditya, et al.
Published: (2024)
A Good Plan is Hard to Find: Aligning Models with Preferences is Misaligned with What Helps Users
by: Balepur, Nishant, et al.
Published: (2025)
by: Balepur, Nishant, et al.
Published: (2025)
What Makes You CLIC: Detection of Croatian Clickbait Headlines
by: Anđelić, Marija, et al.
Published: (2025)
by: Anđelić, Marija, et al.
Published: (2025)
What Makes Diffusion Language Models Super Data Learners?
by: Gao, Zitian, et al.
Published: (2025)
by: Gao, Zitian, et al.
Published: (2025)
UNVEILING: What Makes Linguistics Olympiad Puzzles Tricky for LLMs?
by: Choudhary, Mukund, et al.
Published: (2025)
by: Choudhary, Mukund, et al.
Published: (2025)
Large Language Models Know What Makes Exemplary Contexts
by: Long, Quanyu, et al.
Published: (2024)
by: Long, Quanyu, et al.
Published: (2024)
Collective Constitutional AI: Aligning a Language Model with Public Input
by: Huang, Saffron, et al.
Published: (2024)
by: Huang, Saffron, et al.
Published: (2024)
What Makes In-context Learning Effective for Mathematical Reasoning: A Theoretical Analysis
by: Liu, Jiayu, et al.
Published: (2024)
by: Liu, Jiayu, et al.
Published: (2024)
What Makes LLM Agent Simulations Useful for Policy Practice? An Iterative Design Study in Emergency Preparedness
by: Li, Yuxuan, et al.
Published: (2025)
by: Li, Yuxuan, et al.
Published: (2025)
Similar Items
-
Oolong: Evaluating Long Context Reasoning and Aggregation Capabilities
by: Bertsch, Amanda, et al.
Published: (2025) -
Vocabulary embeddings organize linguistic structure early in language model training
by: Papadimitriou, Isabel, et al.
Published: (2025) -
Investigating the interaction of linguistic and mathematical reasoning in language models using multilingual number puzzles
by: Bhattacharya, Antara Raaghavi, et al.
Published: (2025) -
ReCOGS: How Incidental Details of a Logical Form Overshadow an Evaluation of Semantic Interpretation
by: Wu, Zhengxuan, et al.
Published: (2023) -
What Makes CLIP More Robust to Long-Tailed Pre-Training Data? A Controlled Study for Transferable Insights
by: Wen, Xin, et al.
Published: (2024)