Words in Motion: Extracting Interpretable Control Vectors for Motion Transformers
Fuente:
arXiv
Enregistré dans:
| Auteurs principaux: | Tas, Omer Sahin, Wagner, Royden |
|---|---|
| Format: | Preprint |
| Publié: |
2024
|
| Sujets: | |
| Accès en ligne: | |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
JointMotion: Joint Self-Supervision for Joint Motion Prediction
par: Wagner, Royden, et autres
Publié: (2024)
par: Wagner, Royden, et autres
Publié: (2024)
RedMotion: Motion Prediction via Redundancy Reduction
par: Wagner, Royden, et autres
Publié: (2023)
par: Wagner, Royden, et autres
Publié: (2023)
Divide and Merge: Motion and Semantic Learning in End-to-End Autonomous Driving
par: Shen, Yinzhe, et autres
Publié: (2025)
par: Shen, Yinzhe, et autres
Publié: (2025)
Transformer with Controlled Attention for Synchronous Motion Captioning
par: Radouane, Karim, et autres
Publié: (2024)
par: Radouane, Karim, et autres
Publié: (2024)
SceneMotion: From Agent-Centric Embeddings to Scene-Wide Forecasts
par: Wagner, Royden, et autres
Publié: (2024)
par: Wagner, Royden, et autres
Publié: (2024)
RetroMotion: Retrocausal Motion Forecasting Models are Instructable
par: Wagner, Royden, et autres
Publié: (2025)
par: Wagner, Royden, et autres
Publié: (2025)
Transformer-VQ: Linear-Time Transformers via Vector Quantization
par: Lingle, Lucas D.
Publié: (2023)
par: Lingle, Lucas D.
Publié: (2023)
BAD: Bidirectional Auto-regressive Diffusion for Text-to-Motion Generation
par: Hosseyni, S. Rohollah, et autres
Publié: (2024)
par: Hosseyni, S. Rohollah, et autres
Publié: (2024)
MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks
par: Wu, Yiming, et autres
Publié: (2024)
par: Wu, Yiming, et autres
Publié: (2024)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
par: Kim, Eunji, et autres
Publié: (2024)
par: Kim, Eunji, et autres
Publié: (2024)
MotionStream: Real-Time Video Generation with Interactive Motion Controls
par: Shin, Joonghyuk, et autres
Publié: (2025)
par: Shin, Joonghyuk, et autres
Publié: (2025)
Words That Make Language Models Perceive
par: Wang, Sophie L., et autres
Publié: (2025)
par: Wang, Sophie L., et autres
Publié: (2025)
LMFormer: Lane based Motion Prediction Transformer
par: Yadav, Harsh, et autres
Publié: (2025)
par: Yadav, Harsh, et autres
Publié: (2025)
Towards Understanding Camera Motions in Any Video
par: Lin, Zhiqiu, et autres
Publié: (2025)
par: Lin, Zhiqiu, et autres
Publié: (2025)
Selective, Interpretable, and Motion Consistent Privacy Attribute Obfuscation for Action Recognition
par: Ilic, Filip, et autres
Publié: (2024)
par: Ilic, Filip, et autres
Publié: (2024)
No Training Wheels: Steering Vectors for Bias Correction at Inference Time
par: Gupta, Aviral, et autres
Publié: (2025)
par: Gupta, Aviral, et autres
Publié: (2025)
Interpretable Motion-Attentive Maps: Spatio-Temporally Localizing Concepts in Video Diffusion Transformers
par: Jun, Youngjun, et autres
Publié: (2026)
par: Jun, Youngjun, et autres
Publié: (2026)
Interpretability Needs a New Paradigm
par: Madsen, Andreas, et autres
Publié: (2024)
par: Madsen, Andreas, et autres
Publié: (2024)
Scaling Large Motion Models with Million-Level Human Motions
par: Wang, Ye, et autres
Publié: (2024)
par: Wang, Ye, et autres
Publié: (2024)
Spatio-Temporal Branching for Motion Prediction using Motion Increments
par: Wang, Jiexin, et autres
Publié: (2023)
par: Wang, Jiexin, et autres
Publié: (2023)
CoMotion: Concurrent Multi-person 3D Motion
par: Newell, Alejandro, et autres
Publié: (2025)
par: Newell, Alejandro, et autres
Publié: (2025)
Verbalized Representation Learning for Interpretable Few-Shot Generalization
par: Yang, Cheng-Fu, et autres
Publié: (2024)
par: Yang, Cheng-Fu, et autres
Publié: (2024)
IPO: Interpretable Prompt Optimization for Vision-Language Models
par: Du, Yingjun, et autres
Publié: (2024)
par: Du, Yingjun, et autres
Publié: (2024)
Model Interpretability and Rationale Extraction by Input Mask Optimization
par: Brinner, Marc, et autres
Publié: (2025)
par: Brinner, Marc, et autres
Publié: (2025)
Multimodal Arabic Captioning with Interpretable Visual Concept Integration
par: Elchafei, Passant, et autres
Publié: (2025)
par: Elchafei, Passant, et autres
Publié: (2025)
Textual Steering Vectors Can Improve Visual Understanding in Multimodal Large Language Models
par: Gan, Woody Haosheng, et autres
Publié: (2025)
par: Gan, Woody Haosheng, et autres
Publié: (2025)
From Diffusion to Flow: Efficient Motion Generation in MotionGPT3
par: Ban, Jaymin, et autres
Publié: (2026)
par: Ban, Jaymin, et autres
Publié: (2026)
EchoAgent: Guideline-Centric Reasoning Agent for Echocardiography Measurement and Interpretation
par: Daghyani, Matin, et autres
Publié: (2025)
par: Daghyani, Matin, et autres
Publié: (2025)
General Transform: A Unified Framework for Adaptive Transform to Enhance Representations
par: Budiutama, Gekko, et autres
Publié: (2025)
par: Budiutama, Gekko, et autres
Publié: (2025)
Video Motion Transfer with Diffusion Transformers
par: Pondaven, Alexander, et autres
Publié: (2024)
par: Pondaven, Alexander, et autres
Publié: (2024)
Moto: Latent Motion Token as the Bridging Language for Learning Robot Manipulation from Videos
par: Chen, Yi, et autres
Publié: (2024)
par: Chen, Yi, et autres
Publié: (2024)
StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross Fusion
par: Guo, Ziyu, et autres
Publié: (2025)
par: Guo, Ziyu, et autres
Publié: (2025)
Interpreting and Controlling Model Behavior via Constitutions for Atomic Concept Edits
par: Kalibhat, Neha, et autres
Publié: (2026)
par: Kalibhat, Neha, et autres
Publié: (2026)
A Survey on Transformer Compression
par: Tang, Yehui, et autres
Publié: (2024)
par: Tang, Yehui, et autres
Publié: (2024)
Not All Tokens Matter Equally: Dynamic In-context Vector Distillation with Decisive-Token Supervision for Long-form Medical Report Generation
par: Wu, Ning, et autres
Publié: (2026)
par: Wu, Ning, et autres
Publié: (2026)
CoMusion: Towards Consistent Stochastic Human Motion Prediction via Motion Diffusion
par: Sun, Jiarui, et autres
Publié: (2023)
par: Sun, Jiarui, et autres
Publié: (2023)
Exploring the Limits of Zero Shot Vision Language Models for Hate Meme Detection: The Vulnerabilities and their Interpretations
par: Rizwan, Naquee, et autres
Publié: (2024)
par: Rizwan, Naquee, et autres
Publié: (2024)
CoGen: Learning from Feedback with Coupled Comprehension and Generation
par: Gul, Mustafa Omer, et autres
Publié: (2024)
par: Gul, Mustafa Omer, et autres
Publié: (2024)
MatFormer: Nested Transformer for Elastic Inference
par: Devvrit, et autres
Publié: (2023)
par: Devvrit, et autres
Publié: (2023)
Investigating VLM Hallucination from a Cognitive Psychology Perspective: A First Step Toward Interpretation with Intriguing Observations
par: Liu, Xiangrui, et autres
Publié: (2025)
par: Liu, Xiangrui, et autres
Publié: (2025)
Documents similaires
-
JointMotion: Joint Self-Supervision for Joint Motion Prediction
par: Wagner, Royden, et autres
Publié: (2024) -
RedMotion: Motion Prediction via Redundancy Reduction
par: Wagner, Royden, et autres
Publié: (2023) -
Divide and Merge: Motion and Semantic Learning in End-to-End Autonomous Driving
par: Shen, Yinzhe, et autres
Publié: (2025) -
Transformer with Controlled Attention for Synchronous Motion Captioning
par: Radouane, Karim, et autres
Publié: (2024) -
SceneMotion: From Agent-Centric Embeddings to Scene-Wide Forecasts
par: Wagner, Royden, et autres
Publié: (2024)