Invariant Features in Language Models: Geometric Characterization and Model Attribution
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Dasgupta, Agnibh, Tanvir, Abdullah, Zhong, Xin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Watermarking Language Models through Language Models
von: Dasgupta, Agnibh, et al.
Veröffentlicht: (2024)
von: Dasgupta, Agnibh, et al.
Veröffentlicht: (2024)
TIACam: Text-Anchored Invariant Feature Learning with Auto-Augmentation for Camera-Robust Zero-Watermarking
von: Tanvir, Abdullah All, et al.
Veröffentlicht: (2026)
von: Tanvir, Abdullah All, et al.
Veröffentlicht: (2026)
InvZW: Invariant Feature Learning via Noise-Adversarial Training for Robust Image Zero-Watermarking
von: Tanvir, Abdullah All, et al.
Veröffentlicht: (2025)
von: Tanvir, Abdullah All, et al.
Veröffentlicht: (2025)
Disentangled Representation Learning with Large Language Models for Text-Attributed Graphs
von: Qin, Yijian, et al.
Veröffentlicht: (2023)
von: Qin, Yijian, et al.
Veröffentlicht: (2023)
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)
ReAGent: A Model-agnostic Feature Attribution Method for Generative Language Models
von: Zhao, Zhixue, et al.
Veröffentlicht: (2024)
von: Zhao, Zhixue, et al.
Veröffentlicht: (2024)
Reliable, Adaptable, and Attributable Language Models with Retrieval
von: Asai, Akari, et al.
Veröffentlicht: (2024)
von: Asai, Akari, et al.
Veröffentlicht: (2024)
Don't Ignore the Tail: Decoupling top-K Probabilities for Efficient Language Model Distillation
von: Dasgupta, Sayantan, et al.
Veröffentlicht: (2026)
von: Dasgupta, Sayantan, et al.
Veröffentlicht: (2026)
Neuron-Level Knowledge Attribution in Large Language Models
von: Yu, Zeping, et al.
Veröffentlicht: (2023)
von: Yu, Zeping, et al.
Veröffentlicht: (2023)
Distilling Large Language Models for Text-Attributed Graph Learning
von: Pan, Bo, et al.
Veröffentlicht: (2024)
von: Pan, Bo, et al.
Veröffentlicht: (2024)
Confidence-Modulated Speculative Decoding for Large Language Models
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
von: Sen, Jaydip, et al.
Veröffentlicht: (2025)
Detecting and Characterizing Planning in Language Models
von: Nainani, Jatin, et al.
Veröffentlicht: (2025)
von: Nainani, Jatin, et al.
Veröffentlicht: (2025)
Transferring Linear Features Across Language Models With Model Stitching
von: Chen, Alan, et al.
Veröffentlicht: (2025)
von: Chen, Alan, et al.
Veröffentlicht: (2025)
Learning Multiplex Representations on Text-Attributed Graphs with One Language Model Encoder
von: Jin, Bowen, et al.
Veröffentlicht: (2023)
von: Jin, Bowen, et al.
Veröffentlicht: (2023)
Feature Alignment and Representation Transfer in Knowledge Distillation for Large Language Models
von: Yang, Junjie, et al.
Veröffentlicht: (2025)
von: Yang, Junjie, et al.
Veröffentlicht: (2025)
Rotary Offset Features in Large Language Models
von: Jonasson, André
Veröffentlicht: (2025)
von: Jonasson, André
Veröffentlicht: (2025)
Latent Feature Mining for Predictive Model Enhancement with Large Language Models
von: Li, Bingxuan, et al.
Veröffentlicht: (2024)
von: Li, Bingxuan, et al.
Veröffentlicht: (2024)
JoPA:Explaining Large Language Model's Generation via Joint Prompt Attribution
von: Chang, Yurui, et al.
Veröffentlicht: (2024)
von: Chang, Yurui, et al.
Veröffentlicht: (2024)
Semantic Structure of Feature Space in Large Language Models
von: Kozlowski, Austin C., et al.
Veröffentlicht: (2026)
von: Kozlowski, Austin C., et al.
Veröffentlicht: (2026)
Automatically Interpreting Millions of Features in Large Language Models
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
von: Paulo, Gonçalo, et al.
Veröffentlicht: (2024)
Language Model Training Paradigms for Clinical Feature Embeddings
von: Hu, Yurong, et al.
Veröffentlicht: (2023)
von: Hu, Yurong, et al.
Veröffentlicht: (2023)
ANUBHUTI: A Comprehensive Corpus For Sentiment Analysis In Bangla Regional Languages
von: Kundu, Swastika, et al.
Veröffentlicht: (2025)
von: Kundu, Swastika, et al.
Veröffentlicht: (2025)
MemTrace: Tracing and Attributing Errors in Large Language Model Memory Systems
von: Deng, Xinle, et al.
Veröffentlicht: (2026)
von: Deng, Xinle, et al.
Veröffentlicht: (2026)
Geometric Properties of the Voronoi Tessellation in Latent Semantic Manifolds of Large Language Models
von: Brett, Marshall
Veröffentlicht: (2026)
von: Brett, Marshall
Veröffentlicht: (2026)
ContextCite: Attributing Model Generation to Context
von: Cohen-Wang, Benjamin, et al.
Veröffentlicht: (2024)
von: Cohen-Wang, Benjamin, et al.
Veröffentlicht: (2024)
Persistent Topological Features in Large Language Models
von: Gardinazzi, Yuri, et al.
Veröffentlicht: (2024)
von: Gardinazzi, Yuri, et al.
Veröffentlicht: (2024)
ARGUS: Adaptive Rotation-Invariant Geometric Unsupervised System
von: Sharma, Anantha
Veröffentlicht: (2026)
von: Sharma, Anantha
Veröffentlicht: (2026)
Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
von: Li, Michael, et al.
Veröffentlicht: (2025)
von: Li, Michael, et al.
Veröffentlicht: (2025)
Analyze Feature Flow to Enhance Interpretation and Steering in Language Models
von: Laptev, Daniil, et al.
Veröffentlicht: (2025)
von: Laptev, Daniil, et al.
Veröffentlicht: (2025)
Multi-Attribute Steering of Language Models via Targeted Intervention
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
von: Nguyen, Duy, et al.
Veröffentlicht: (2025)
HyperBERT: Mixing Hypergraph-Aware Layers with Language Models for Node Classification on Text-Attributed Hypergraphs
von: Bazaga, Adrián, et al.
Veröffentlicht: (2024)
von: Bazaga, Adrián, et al.
Veröffentlicht: (2024)
Large Language Models are Null-Shot Learners
von: Taveekitworachai, Pittawat, et al.
Veröffentlicht: (2024)
von: Taveekitworachai, Pittawat, et al.
Veröffentlicht: (2024)
Harnessing Large Language Models as Post-hoc Correctors
von: Zhong, Zhiqiang, et al.
Veröffentlicht: (2024)
von: Zhong, Zhiqiang, et al.
Veröffentlicht: (2024)
How do Large Language Models Navigate Conflicts between Honesty and Helpfulness?
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
von: Liu, Ryan, et al.
Veröffentlicht: (2024)
Feature Extraction and Steering for Enhanced Chain-of-Thought Reasoning in Language Models
von: Li, Zihao, et al.
Veröffentlicht: (2025)
von: Li, Zihao, et al.
Veröffentlicht: (2025)
On the Geometric Structure of Layer Updates in Deep Language Models
von: Yoo, Jun-Sik
Veröffentlicht: (2026)
von: Yoo, Jun-Sik
Veröffentlicht: (2026)
Enhancing Authorship Attribution through Embedding Fusion: A Novel Approach with Masked and Encoder-Decoder Language Models
von: Kaushik, Arjun Ramesh, et al.
Veröffentlicht: (2024)
von: Kaushik, Arjun Ramesh, et al.
Veröffentlicht: (2024)
Cite Pretrain: Retrieval-Free Knowledge Attribution for Large Language Models
von: Huang, Yukun, et al.
Veröffentlicht: (2025)
von: Huang, Yukun, et al.
Veröffentlicht: (2025)
Large Language Models as Topological Structure Enhancers for Text-Attributed Graphs
von: Sun, Shengyin, et al.
Veröffentlicht: (2023)
von: Sun, Shengyin, et al.
Veröffentlicht: (2023)
Can Authorship Attribution Models Distinguish Speakers in Speech Transcripts?
von: Aggazzotti, Cristina, et al.
Veröffentlicht: (2023)
von: Aggazzotti, Cristina, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
Watermarking Language Models through Language Models
von: Dasgupta, Agnibh, et al.
Veröffentlicht: (2024) -
TIACam: Text-Anchored Invariant Feature Learning with Auto-Augmentation for Camera-Robust Zero-Watermarking
von: Tanvir, Abdullah All, et al.
Veröffentlicht: (2026) -
InvZW: Invariant Feature Learning via Noise-Adversarial Training for Robust Image Zero-Watermarking
von: Tanvir, Abdullah All, et al.
Veröffentlicht: (2025) -
Disentangled Representation Learning with Large Language Models for Text-Attributed Graphs
von: Qin, Yijian, et al.
Veröffentlicht: (2023) -
Beyond Linear Steering: Unified Multi-Attribute Control for Language Models
von: Oozeer, Narmeen, et al.
Veröffentlicht: (2025)