Salvato in:
| Autori principali: | Khona, Mikail, Okawa, Maya, Hula, Jan, Ramesh, Rahul, Nishi, Kento, Dick, Robert, Lubana, Ekdeep Singh, Tanaka, Hidenori |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2402.07757 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing
di: Nishi, Kento, et al.
Pubblicazione: (2024)
di: Nishi, Kento, et al.
Pubblicazione: (2024)
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks
di: Ramesh, Rahul, et al.
Pubblicazione: (2023)
di: Ramesh, Rahul, et al.
Pubblicazione: (2023)
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
di: Okawa, Maya, et al.
Pubblicazione: (2023)
di: Okawa, Maya, et al.
Pubblicazione: (2023)
ICLR: In-Context Learning of Representations
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
A Percolation Model of Emergence: Analyzing Transformers Trained on a Formal Language
di: Lubana, Ekdeep Singh, et al.
Pubblicazione: (2024)
di: Lubana, Ekdeep Singh, et al.
Pubblicazione: (2024)
Swing-by Dynamics in Concept Learning and Compositional Generalization
di: Yang, Yongyi, et al.
Pubblicazione: (2024)
di: Yang, Yongyi, et al.
Pubblicazione: (2024)
Emergence of Hierarchical Emotion Organization in Large Language Models
di: Zhao, Bo, et al.
Pubblicazione: (2025)
di: Zhao, Bo, et al.
Pubblicazione: (2025)
In-Context Learning Dynamics with Random Binary Sequences
di: Bigelow, Eric J., et al.
Pubblicazione: (2023)
di: Bigelow, Eric J., et al.
Pubblicazione: (2023)
Competition Dynamics Shape Algorithmic Phases of In-Context Learning
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
di: Park, Core Francisco, et al.
Pubblicazione: (2024)
Abrupt Learning in Transformers: A Case Study on Matrix Completion
di: Gopalani, Pulkit, et al.
Pubblicazione: (2024)
di: Gopalani, Pulkit, et al.
Pubblicazione: (2024)
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
di: Jain, Samyak, et al.
Pubblicazione: (2023)
di: Jain, Samyak, et al.
Pubblicazione: (2023)
In-Context Learning Strategies Emerge Rationally
di: Wurgaft, Daniel, et al.
Pubblicazione: (2025)
di: Wurgaft, Daniel, et al.
Pubblicazione: (2025)
Towards Reliable Evaluation of Behavior Steering Interventions in LLMs
di: Pres, Itamar, et al.
Pubblicazione: (2024)
di: Pres, Itamar, et al.
Pubblicazione: (2024)
How Do LLMs Persuade? Linear Probes Can Uncover Persuasion Dynamics in Multi-Turn Conversations
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2025)
di: Jaipersaud, Brandon, et al.
Pubblicazione: (2025)
Belief Dynamics Reveal the Dual Nature of In-Context Learning and Activation Steering
di: Bigelow, Eric, et al.
Pubblicazione: (2025)
di: Bigelow, Eric, et al.
Pubblicazione: (2025)
In-Context Learning of Energy Functions
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)
Are language models aware of the road not taken? Token-level uncertainty and hidden state dynamics
di: Zur, Amir, et al.
Pubblicazione: (2025)
di: Zur, Amir, et al.
Pubblicazione: (2025)
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
di: Bohacek, Matyas, et al.
Pubblicazione: (2025)
di: Bohacek, Matyas, et al.
Pubblicazione: (2025)
Analyzing (In)Abilities of SAEs via Formal Languages
di: Menon, Abhinav, et al.
Pubblicazione: (2024)
di: Menon, Abhinav, et al.
Pubblicazione: (2024)
Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry
di: Hindupur, Sai Sumedh R., et al.
Pubblicazione: (2025)
di: Hindupur, Sai Sumedh R., et al.
Pubblicazione: (2025)
From Flat to Hierarchical: Extracting Sparse Representations with Matching Pursuit
di: Costa, Valérie, et al.
Pubblicazione: (2025)
di: Costa, Valérie, et al.
Pubblicazione: (2025)
Evaluating Sparse Autoencoders: From Shallow Design to Matching Pursuit
di: Costa, Valérie, et al.
Pubblicazione: (2025)
di: Costa, Valérie, et al.
Pubblicazione: (2025)
Understanding GNNs for Boolean Satisfiability through Approximation Algorithms
di: Hůla, Jan, et al.
Pubblicazione: (2024)
di: Hůla, Jan, et al.
Pubblicazione: (2024)
Uncovering Latent Memories: Assessing Data Leakage and Memorization Patterns in Frontier AI Models
di: Duan, Sunny, et al.
Pubblicazione: (2024)
di: Duan, Sunny, et al.
Pubblicazione: (2024)
The Impact of Off-Policy Training Data on Probe Generalisation
di: Kirch, Nathalie, et al.
Pubblicazione: (2025)
di: Kirch, Nathalie, et al.
Pubblicazione: (2025)
From Isolation to Entanglement: When Do Interpretability Methods Identify and Disentangle Known Concepts?
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
di: Mueller, Aaron, et al.
Pubblicazione: (2025)
Behind India's ChatGPT Conversations: A Retrospective Analysis of 238 Unedited User Prompts
di: Khona, Kalyani
Pubblicazione: (2025)
di: Khona, Kalyani
Pubblicazione: (2025)
Detecting High-Stakes Interactions with Activation Probes
di: McKenzie, Alex, et al.
Pubblicazione: (2025)
di: McKenzie, Alex, et al.
Pubblicazione: (2025)
Towards an Improved Understanding and Utilization of Maximum Manifold Capacity Representations
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)
di: Schaeffer, Rylan, et al.
Pubblicazione: (2024)
What Makes and Breaks Safety Fine-tuning? A Mechanistic Study
di: Jain, Samyak, et al.
Pubblicazione: (2024)
di: Jain, Samyak, et al.
Pubblicazione: (2024)
Features as Rewards: Scalable Supervision for Open-Ended Tasks via Interpretability
di: Prasad, Aaditya Vikram, et al.
Pubblicazione: (2026)
di: Prasad, Aaditya Vikram, et al.
Pubblicazione: (2026)
Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space
di: Bigelow, Eric, et al.
Pubblicazione: (2026)
di: Bigelow, Eric, et al.
Pubblicazione: (2026)
When Is Collective Intelligence a Lottery? Multi-Agent Scaling Laws for Memetic Drift in LLMs
di: Tanaka, Hidenori
Pubblicazione: (2026)
di: Tanaka, Hidenori
Pubblicazione: (2026)
Morphometric analysis of turfgrass using digital three‐dimensional technology and its application in breeding
di: Hidenori Tanaka
Pubblicazione: (2025)
di: Hidenori Tanaka
Pubblicazione: (2025)
New Constraints on Gauged U(1)$_{L_μ-L_τ}$ Models via $Z-Z'$ Mixing
di: Asai, Kento, et al.
Pubblicazione: (2024)
di: Asai, Kento, et al.
Pubblicazione: (2024)
Minimizing the Weighted Number of Tardy Jobs: Data-Driven Heuristic for Single-Machine Scheduling
di: Antonov, Nikolai, et al.
Pubblicazione: (2025)
di: Antonov, Nikolai, et al.
Pubblicazione: (2025)
Aggregated Multi-output Gaussian Processes with Knowledge Transfer Across Domains
di: Tanaka, Yusuke, et al.
Pubblicazione: (2022)
di: Tanaka, Yusuke, et al.
Pubblicazione: (2022)
Enhancing Agentic Textual Graph Retrieval with Synthetic Stepwise Supervision
di: Chang, Ge, et al.
Pubblicazione: (2025)
di: Chang, Ge, et al.
Pubblicazione: (2025)
The Shape of Beliefs: Geometry, Dynamics, and Interventions along Representation Manifolds of Language Models' Posteriors
di: Sarfati, Raphaël, et al.
Pubblicazione: (2026)
di: Sarfati, Raphaël, et al.
Pubblicazione: (2026)
Documenti analoghi
-
Representation Shattering in Transformers: A Synthetic Study with Knowledge Editing
di: Nishi, Kento, et al.
Pubblicazione: (2024) -
Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks
di: Ramesh, Rahul, et al.
Pubblicazione: (2023) -
Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task
di: Okawa, Maya, et al.
Pubblicazione: (2023) -
ICLR: In-Context Learning of Representations
di: Park, Core Francisco, et al.
Pubblicazione: (2024) -
Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept Space
di: Park, Core Francisco, et al.
Pubblicazione: (2024)