X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization
Fuente:
arXiv
Saved in:
| Main Authors: | Kukleva, Anna, Sener, Fadime, Remelli, Edoardo, Tekin, Bugra, Sauser, Eric, Schiele, Bernt, Ma, Shugao |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
DiffH2O: Diffusion-Based Synthesis of Hand-Object Interactions from Textual Descriptions
by: Christen, Sammy, et al.
Published: (2024)
by: Christen, Sammy, et al.
Published: (2024)
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
by: Chatterjee, Dibyadip, et al.
Published: (2025)
by: Chatterjee, Dibyadip, et al.
Published: (2025)
PALM: A Dataset and Baseline for Learning Multi-subject Hand Prior
by: Fan, Zicong, et al.
Published: (2025)
by: Fan, Zicong, et al.
Published: (2025)
OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning
by: Ahmed, Noor, et al.
Published: (2024)
by: Ahmed, Noor, et al.
Published: (2024)
Do Instance Priors Help Weakly Supervised Semantic Segmentation?
by: Das, Anurag, et al.
Published: (2026)
by: Das, Anurag, et al.
Published: (2026)
On the Utility of 3D Hand Poses for Action Recognition
by: Shamil, Md Salman, et al.
Published: (2024)
by: Shamil, Md Salman, et al.
Published: (2024)
MM-TS: Multi-Modal Temperature and Margin Schedules for Contrastive Learning with Long-Tail Data
by: Sheludzko, Siarhei, et al.
Published: (2026)
by: Sheludzko, Siarhei, et al.
Published: (2026)
Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs
by: Kuzucu, Selim, et al.
Published: (2025)
by: Kuzucu, Selim, et al.
Published: (2025)
Context-Enhanced Memory-Refined Transformer for Online Action Detection
by: Pang, Zhanzhong, et al.
Published: (2025)
by: Pang, Zhanzhong, et al.
Published: (2025)
HowToCaption: Prompting LLMs to Transform Video Annotations at Scale
by: Shvetsova, Nina, et al.
Published: (2023)
by: Shvetsova, Nina, et al.
Published: (2023)
On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding
by: Pang, Zhanzhong, et al.
Published: (2026)
by: Pang, Zhanzhong, et al.
Published: (2026)
HuMoCon: Concept Discovery for Human Motion Understanding
by: Fang, Qihang, et al.
Published: (2025)
by: Fang, Qihang, et al.
Published: (2025)
RefAM: Attention Magnets for Zero-Shot Referral Segmentation
by: Kukleva, Anna, et al.
Published: (2025)
by: Kukleva, Anna, et al.
Published: (2025)
Cost-Sensitive Learning for Long-Tailed Temporal Action Segmentation
by: Pang, Zhanzhong, et al.
Published: (2025)
by: Pang, Zhanzhong, et al.
Published: (2025)
Long-Tail Temporal Action Segmentation with Group-wise Temporal Logit Adjustment
by: Pang, Zhanzhong, et al.
Published: (2024)
by: Pang, Zhanzhong, et al.
Published: (2024)
PersonaHOI: Effortlessly Improving Personalized Face with Human-Object Interaction Generation
by: Hu, Xinting, et al.
Published: (2025)
by: Hu, Xinting, et al.
Published: (2025)
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
by: Pang, Zhanzhong, et al.
Published: (2026)
by: Pang, Zhanzhong, et al.
Published: (2026)
Temporal Concept Dynamics in Diffusion Models via Prompt-Conditioned Interventions
by: Gorgun, Ada, et al.
Published: (2025)
by: Gorgun, Ada, et al.
Published: (2025)
DWDN: Deep Wiener Deconvolution Network for Non-Blind Image Deblurring
by: Dong, Jiangxin, et al.
Published: (2021)
by: Dong, Jiangxin, et al.
Published: (2021)
VITAL: More Understandable Feature Visualization through Distribution Alignment and Relevant Information Flow
by: Gorgun, Ada, et al.
Published: (2025)
by: Gorgun, Ada, et al.
Published: (2025)
SimNP: Learning Self-Similarity Priors Between Neural Points
by: Wewer, Christopher, et al.
Published: (2023)
by: Wewer, Christopher, et al.
Published: (2023)
Sp2360: Sparse-view 360 Scene Reconstruction using Cascaded 2D Diffusion Priors
by: Paul, Soumava, et al.
Published: (2024)
by: Paul, Soumava, et al.
Published: (2024)
Don't Pause! Every prediction matters in a streaming video
by: Chatterjee, Dibyadip, et al.
Published: (2026)
by: Chatterjee, Dibyadip, et al.
Published: (2026)
Optimising for Interpretability: Convolutional Dynamic Alignment Networks
by: Böhle, Moritz, et al.
Published: (2021)
by: Böhle, Moritz, et al.
Published: (2021)
Towards Better Understanding Attribution Methods
by: Rao, Sukrut, et al.
Published: (2022)
by: Rao, Sukrut, et al.
Published: (2022)
Spatial Reasoners for Continuous Variables in Any Domain
by: Pogodzinski, Bart, et al.
Published: (2025)
by: Pogodzinski, Bart, et al.
Published: (2025)
Spatial Reasoning with Denoising Models
by: Wewer, Christopher, et al.
Published: (2025)
by: Wewer, Christopher, et al.
Published: (2025)
Rewis3d: Reconstruction Improves Weakly-Supervised Semantic Segmentation
by: Ernst, Jonas, et al.
Published: (2026)
by: Ernst, Jonas, et al.
Published: (2026)
Scribbles for All: Benchmarking Scribble Supervised Segmentation Across Datasets
by: Boettcher, Wolfgang, et al.
Published: (2024)
by: Boettcher, Wolfgang, et al.
Published: (2024)
latentSplat: Autoencoding Variational Gaussians for Fast Generalizable 3D Reconstruction
by: Wewer, Christopher, et al.
Published: (2024)
by: Wewer, Christopher, et al.
Published: (2024)
Better Understanding Differences in Attribution Methods via Systematic Evaluations
by: Rao, Sukrut, et al.
Published: (2023)
by: Rao, Sukrut, et al.
Published: (2023)
Adversarial Training against Location-Optimized Adversarial Patches
by: Rao, Sukrut, et al.
Published: (2020)
by: Rao, Sukrut, et al.
Published: (2020)
CigTime: Corrective Instruction Generation Through Inverse Motion Editing
by: Fang, Qihang, et al.
Published: (2024)
by: Fang, Qihang, et al.
Published: (2024)
ClipTTT: CLIP-Guided Test-Time Training Helps LVLMs See Better
by: Nath, Mriganka, et al.
Published: (2026)
by: Nath, Mriganka, et al.
Published: (2026)
B-cos Alignment for Inherently Interpretable CNNs and Vision Transformers
by: Böhle, Moritz, et al.
Published: (2023)
by: Böhle, Moritz, et al.
Published: (2023)
AIM: Amending Inherent Interpretability via Self-Supervised Masking
by: Alshami, Eyad, et al.
Published: (2025)
by: Alshami, Eyad, et al.
Published: (2025)
How to Probe: Simple Yet Effective Techniques for Improving Post-hoc Explanations
by: Gairola, Siddhartha, et al.
Published: (2025)
by: Gairola, Siddhartha, et al.
Published: (2025)
MTR++: Multi-Agent Motion Prediction with Symmetric Scene Modeling and Guided Intention Querying
by: Shi, Shaoshuai, et al.
Published: (2023)
by: Shi, Shaoshuai, et al.
Published: (2023)
MTA-CLIP: Language-Guided Semantic Segmentation with Mask-Text Alignment
by: Das, Anurag, et al.
Published: (2024)
by: Das, Anurag, et al.
Published: (2024)
MEt3R: Measuring Multi-View Consistency in Generated Images
by: Asim, Mohammad, et al.
Published: (2025)
by: Asim, Mohammad, et al.
Published: (2025)
Similar Items
-
DiffH2O: Diffusion-Based Synthesis of Hand-Object Interactions from Textual Descriptions
by: Christen, Sammy, et al.
Published: (2024) -
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
by: Chatterjee, Dibyadip, et al.
Published: (2025) -
PALM: A Dataset and Baseline for Learning Multi-subject Hand Prior
by: Fan, Zicong, et al.
Published: (2025) -
OrCo: Towards Better Generalization via Orthogonality and Contrast for Few-Shot Class-Incremental Learning
by: Ahmed, Noor, et al.
Published: (2024) -
Do Instance Priors Help Weakly Supervised Semantic Segmentation?
by: Das, Anurag, et al.
Published: (2026)