Gespeichert in:
| Hauptverfasser: | Levy, Ido, Paradise, Orr, Carmeli, Boaz, Meir, Ron, Goldwasser, Shafi, Belinkov, Yonatan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2502.07552 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Investigating the Development of Task-Oriented Communication in Vision-Language Models
von: Carmeli, Boaz, et al.
Veröffentlicht: (2026)
von: Carmeli, Boaz, et al.
Veröffentlicht: (2026)
Concept-Best-Matching: Evaluating Compositionality in Emergent Communication
von: Carmeli, Boaz, et al.
Veröffentlicht: (2024)
von: Carmeli, Boaz, et al.
Veröffentlicht: (2024)
CtD: Composition through Decomposition in Emergent Communication
von: Carmeli, Boaz, et al.
Veröffentlicht: (2026)
von: Carmeli, Boaz, et al.
Veröffentlicht: (2026)
Semantics and Spatiality of Emergent Communication
von: Zion, Rotem Ben, et al.
Veröffentlicht: (2024)
von: Zion, Rotem Ben, et al.
Veröffentlicht: (2024)
Will it Merge? On The Causes of Model Mergeability
von: Rahamim, Adir, et al.
Veröffentlicht: (2026)
von: Rahamim, Adir, et al.
Veröffentlicht: (2026)
SAEs Are Good for Steering -- If You Select the Right Features
von: Arad, Dana, et al.
Veröffentlicht: (2025)
von: Arad, Dana, et al.
Veröffentlicht: (2025)
Planted in Pretraining, Swayed by Finetuning: A Case Study on the Origins of Cognitive Biases in LLMs
von: Itzhak, Itay, et al.
Veröffentlicht: (2025)
von: Itzhak, Itay, et al.
Veröffentlicht: (2025)
Models That Prove Their Own Correctness
von: Amit, Noga, et al.
Veröffentlicht: (2024)
von: Amit, Noga, et al.
Veröffentlicht: (2024)
Backward Lens: Projecting Language Model Gradients into the Vocabulary Space
von: Katz, Shahar, et al.
Veröffentlicht: (2024)
von: Katz, Shahar, et al.
Veröffentlicht: (2024)
From Feelings to Metrics: Understanding and Formalizing How Users Vibe-Test LLMs
von: Itzhak, Itay, et al.
Veröffentlicht: (2026)
von: Itzhak, Itay, et al.
Veröffentlicht: (2026)
Answer, Assemble, Ace: Understanding How LMs Answer Multiple Choice Questions
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
von: Wiegreffe, Sarah, et al.
Veröffentlicht: (2024)
DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
von: Ventura, Mor, et al.
Veröffentlicht: (2025)
von: Ventura, Mor, et al.
Veröffentlicht: (2025)
On Non-interactive Evaluation of Animal Communication Translators
von: Paradise, Orr, et al.
Veröffentlicht: (2025)
von: Paradise, Orr, et al.
Veröffentlicht: (2025)
Learning Randomized Reductions
von: Erata, Ferhat, et al.
Veröffentlicht: (2024)
von: Erata, Ferhat, et al.
Veröffentlicht: (2024)
LLM-Human Pipeline for Cultural Context Grounding of Conversations
von: Pujari, Rajkumar, et al.
Veröffentlicht: (2024)
von: Pujari, Rajkumar, et al.
Veröffentlicht: (2024)
Instructed to Bias: Instruction-Tuned Language Models Exhibit Emergent Cognitive Bias
von: Itzhak, Itay, et al.
Veröffentlicht: (2023)
von: Itzhak, Itay, et al.
Veröffentlicht: (2023)
Old Habits Die Hard: How Conversational History Geometrically Traps LLMs
von: Simhi, Adi, et al.
Veröffentlicht: (2026)
von: Simhi, Adi, et al.
Veröffentlicht: (2026)
ParallelPARC: A Scalable Pipeline for Generating Natural-Language Analogies
von: Sultan, Oren, et al.
Veröffentlicht: (2024)
von: Sultan, Oren, et al.
Veröffentlicht: (2024)
BlackboxNLP-2025 MIB Shared Task: Improving Circuit Faithfulness via Better Edge Selection
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2025)
von: Nikankin, Yaniv, et al.
Veröffentlicht: (2025)
Emergent Communication Pretraining for Few-Shot Machine Translation
von: Li, Yaoyiran, et al.
Veröffentlicht: (2020)
von: Li, Yaoyiran, et al.
Veröffentlicht: (2020)
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models
von: Marks, Samuel, et al.
Veröffentlicht: (2024)
von: Marks, Samuel, et al.
Veröffentlicht: (2024)
Genie: Achieving Human Parity in Content-Grounded Datasets Generation
von: Yehudai, Asaf, et al.
Veröffentlicht: (2024)
von: Yehudai, Asaf, et al.
Veröffentlicht: (2024)
Findings of the BlackboxNLP 2025 Shared Task: Localizing Circuits and Causal Variables in Language Models
von: Arad, Dana, et al.
Veröffentlicht: (2025)
von: Arad, Dana, et al.
Veröffentlicht: (2025)
CoLa: Learning to Interactively Collaborate with Large Language Models
von: Sharma, Abhishek, et al.
Veröffentlicht: (2025)
von: Sharma, Abhishek, et al.
Veröffentlicht: (2025)
Splits! Flexible Sociocultural Linguistic Investigation at Scale
von: Caplan, Eylon, et al.
Veröffentlicht: (2025)
von: Caplan, Eylon, et al.
Veröffentlicht: (2025)
Confidence Regulation Neurons in Language Models
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
von: Stolfo, Alessandro, et al.
Veröffentlicht: (2024)
EmoGist: Efficient In-Context Learning for Visual Emotion Understanding
von: Seoh, Ronald, et al.
Veröffentlicht: (2025)
von: Seoh, Ronald, et al.
Veröffentlicht: (2025)
Post-hoc Study of Climate Microtargeting on Social Media Ads with LLMs: Thematic Insights and Fairness Evaluation
von: Islam, Tunazzina, et al.
Veröffentlicht: (2024)
von: Islam, Tunazzina, et al.
Veröffentlicht: (2024)
Generating Benchmarks for Factuality Evaluation of Language Models
von: Muhlgay, Dor, et al.
Veröffentlicht: (2023)
von: Muhlgay, Dor, et al.
Veröffentlicht: (2023)
A Cryptographic Perspective on Mitigation vs. Detection in Machine Learning
von: Gluch, Greg, et al.
Veröffentlicht: (2025)
von: Gluch, Greg, et al.
Veröffentlicht: (2025)
Unsupervised Representation Learning - an Invariant Risk Minimization Perspective
von: Norman, Yotam, et al.
Veröffentlicht: (2025)
von: Norman, Yotam, et al.
Veröffentlicht: (2025)
Can LLMs Assist Annotators in Identifying Morality Frames? -- Case Study on Vaccination Debate on Social Media
von: Islam, Tunazzina, et al.
Veröffentlicht: (2025)
von: Islam, Tunazzina, et al.
Veröffentlicht: (2025)
Discovering Latent Themes in Social Media Messaging: A Machine-in-the-Loop Approach Integrating LLMs
von: Islam, Tunazzina, et al.
Veröffentlicht: (2024)
von: Islam, Tunazzina, et al.
Veröffentlicht: (2024)
Uncovering Latent Arguments in Social Media Messaging by Employing LLMs-in-the-Loop Strategy
von: Islam, Tunazzina, et al.
Veröffentlicht: (2024)
von: Islam, Tunazzina, et al.
Veröffentlicht: (2024)
LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
von: Orgad, Hadas, et al.
Veröffentlicht: (2024)
Large Language Models Generate Harmful Content Using a Distinct, Unified Mechanism
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)
von: Orgad, Hadas, et al.
Veröffentlicht: (2026)
VideoAgent: Long-form Video Understanding with Large Language Model as Agent
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
von: Wang, Xiaohan, et al.
Veröffentlicht: (2024)
ContraSim -- Analyzing Neural Representations Based on Contrastive Learning
von: Rahamim, Adir, et al.
Veröffentlicht: (2023)
von: Rahamim, Adir, et al.
Veröffentlicht: (2023)
Bridging the Gap in Bangla Healthcare: Machine Learning Based Disease Prediction Using a Symptoms-Disease Dataset
von: Zannat, Rowzatul, et al.
Veröffentlicht: (2026)
von: Zannat, Rowzatul, et al.
Veröffentlicht: (2026)
Truth is Universal: Robust Detection of Lies in LLMs
von: Bürger, Lennart, et al.
Veröffentlicht: (2024)
von: Bürger, Lennart, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Investigating the Development of Task-Oriented Communication in Vision-Language Models
von: Carmeli, Boaz, et al.
Veröffentlicht: (2026) -
Concept-Best-Matching: Evaluating Compositionality in Emergent Communication
von: Carmeli, Boaz, et al.
Veröffentlicht: (2024) -
CtD: Composition through Decomposition in Emergent Communication
von: Carmeli, Boaz, et al.
Veröffentlicht: (2026) -
Semantics and Spatiality of Emergent Communication
von: Zion, Rotem Ben, et al.
Veröffentlicht: (2024) -
Will it Merge? On The Causes of Model Mergeability
von: Rahamim, Adir, et al.
Veröffentlicht: (2026)