Gespeichert in:
| Hauptverfasser: | Cai, Yichao, Liu, Yuhang, Gao, Erdun, Jiang, Tianjiao, Zhang, Zhen, Hengel, Anton van den, Shi, Javen Qinfeng |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2504.10143 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning
von: Jiang, Tianjiao, et al.
Veröffentlicht: (2025)
von: Jiang, Tianjiao, et al.
Veröffentlicht: (2025)
Beyond DAGs: A Latent Partial Causal Model for Multimodal Learning
von: Liu, Yuhang, et al.
Veröffentlicht: (2024)
von: Liu, Yuhang, et al.
Veröffentlicht: (2024)
CLAP: Isolating Content from Style through Contrastive Learning with Augmented Prompts
von: Cai, Yichao, et al.
Veröffentlicht: (2023)
von: Cai, Yichao, et al.
Veröffentlicht: (2023)
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
von: Tian, Mingkai, et al.
Veröffentlicht: (2025)
von: Tian, Mingkai, et al.
Veröffentlicht: (2025)
Concept Component Analysis: A Principled Approach for Concept Extraction in LLMs
von: Liu, Yuhang, et al.
Veröffentlicht: (2026)
von: Liu, Yuhang, et al.
Veröffentlicht: (2026)
Learning to Reason and Navigate: Parameter Efficient Action Planning with Large Language Models
von: Mohammadi, Bahram, et al.
Veröffentlicht: (2025)
von: Mohammadi, Bahram, et al.
Veröffentlicht: (2025)
I Predict Therefore I Am: Is Next Token Prediction Enough to Learn Human-Interpretable Concepts from Data?
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
von: Liu, Yuhang, et al.
Veröffentlicht: (2025)
The Geometric Mechanics of Contrastive Representation Learning: Alignment Potentials, Entropic Dispersion, and Cross-modal Divergence
von: Cai, Yichao, et al.
Veröffentlicht: (2026)
von: Cai, Yichao, et al.
Veröffentlicht: (2026)
A Survey on Deep Neural Network Pruning-Taxonomy, Comparison, Analysis, and Recommendations
von: Cheng, Hongrong, et al.
Veröffentlicht: (2023)
von: Cheng, Hongrong, et al.
Veröffentlicht: (2023)
A Simple-but-effective Baseline for Training-free Class-Agnostic Counting
von: Lin, Yuhao, et al.
Veröffentlicht: (2024)
von: Lin, Yuhao, et al.
Veröffentlicht: (2024)
Can You Learn to See Without Images? Procedural Warm-Up for Vision Transformers
von: Shinnick, Zachary, et al.
Veröffentlicht: (2025)
von: Shinnick, Zachary, et al.
Veröffentlicht: (2025)
Towards Identifiable Latent Additive Noise Models
von: Liu, Yuhang, et al.
Veröffentlicht: (2024)
von: Liu, Yuhang, et al.
Veröffentlicht: (2024)
What Makes a Representation Good for Single-Cell Perturbation Prediction?
von: Jiang, Wenkang, et al.
Veröffentlicht: (2026)
von: Jiang, Wenkang, et al.
Veröffentlicht: (2026)
Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors
von: Xia, Jiatong, et al.
Veröffentlicht: (2026)
von: Xia, Jiatong, et al.
Veröffentlicht: (2026)
Premonition: Using Generative Models to Preempt Future Data Changes in Continual Learning
von: McDonnell, Mark D., et al.
Veröffentlicht: (2024)
von: McDonnell, Mark D., et al.
Veröffentlicht: (2024)
Scaling up Multi-domain Semantic Segmentation with Sentence Embeddings
von: Yin, Wei, et al.
Veröffentlicht: (2022)
von: Yin, Wei, et al.
Veröffentlicht: (2022)
Augmented Commonsense Knowledge for Remote Object Grounding
von: Mohammadi, Bahram, et al.
Veröffentlicht: (2024)
von: Mohammadi, Bahram, et al.
Veröffentlicht: (2024)
Let Your Video Listen to Your Music!
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
von: Zhang, Xinyu, et al.
Veröffentlicht: (2025)
Frame-wise Conditioning Adaptation for Fine-Tuning Diffusion Models in Text-to-Video Prediction
von: Liu, Zheyuan, et al.
Veröffentlicht: (2025)
von: Liu, Zheyuan, et al.
Veröffentlicht: (2025)
RanPAC: Random Projections and Pre-trained Models for Continual Learning
von: McDonnell, Mark D., et al.
Veröffentlicht: (2023)
von: McDonnell, Mark D., et al.
Veröffentlicht: (2023)
Continual Learning on CLIP via Incremental Prompt Tuning with Intrinsic Textual Anchors
von: Lu, Haodong, et al.
Veröffentlicht: (2025)
von: Lu, Haodong, et al.
Veröffentlicht: (2025)
Chem4DLLM: 4D Multimodal LLMs for Chemical Dynamics Understanding
von: Li, Xinyu, et al.
Veröffentlicht: (2026)
von: Li, Xinyu, et al.
Veröffentlicht: (2026)
Hierarchical Process Reward Models are Symbolic Vision Learners
von: Zhang, Shan, et al.
Veröffentlicht: (2025)
von: Zhang, Shan, et al.
Veröffentlicht: (2025)
ViewFusion: Towards Multi-View Consistency via Interpolated Denoising
von: Yang, Xianghui, et al.
Veröffentlicht: (2024)
von: Yang, Xianghui, et al.
Veröffentlicht: (2024)
InvariantStock: Learning Invariant Features for Mastering the Shifting Market
von: Cao, Haiyao, et al.
Veröffentlicht: (2024)
von: Cao, Haiyao, et al.
Veröffentlicht: (2024)
Learning Latent Dynamical Causal Processes for Single-Cell Perturbation Prediction
von: Jiang, Wenkang, et al.
Veröffentlicht: (2026)
von: Jiang, Wenkang, et al.
Veröffentlicht: (2026)
Knowledge Composition using Task Vectors with Learned Anisotropic Scaling
von: Zhang, Frederic Z., et al.
Veröffentlicht: (2024)
von: Zhang, Frederic Z., et al.
Veröffentlicht: (2024)
Source-Free Unsupervised Domain Adaptation with Hypothesis Consolidation of Prediction Rationale
von: Shu, Yangyang, et al.
Veröffentlicht: (2024)
von: Shu, Yangyang, et al.
Veröffentlicht: (2024)
Categorical Keypoint Positional Embedding for Robust Animal Re-Identification
von: Lin, Yuhao, et al.
Veröffentlicht: (2024)
von: Lin, Yuhao, et al.
Veröffentlicht: (2024)
Representation Space Constrained Learning with Modality Decoupling for Multimodal Object Detection
von: Shao, YiKang, et al.
Veröffentlicht: (2025)
von: Shao, YiKang, et al.
Veröffentlicht: (2025)
Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment
von: Xu, Rui, et al.
Veröffentlicht: (2025)
von: Xu, Rui, et al.
Veröffentlicht: (2025)
DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
von: Qian, Chengxuan, et al.
Veröffentlicht: (2025)
von: Qian, Chengxuan, et al.
Veröffentlicht: (2025)
Towards Higher Effective Rank in Parameter-efficient Fine-tuning using Khatri--Rao Product
von: Albert, Paul, et al.
Veröffentlicht: (2025)
von: Albert, Paul, et al.
Veröffentlicht: (2025)
Math Blind: Failures in Diagram Understanding Undermine Reasoning in MLLMs
von: Sun, Yanpeng, et al.
Veröffentlicht: (2025)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2025)
ProgRoCC: A Progressive Approach to Rough Crowd Counting
von: Jiang, Shengqin, et al.
Veröffentlicht: (2025)
von: Jiang, Shengqin, et al.
Veröffentlicht: (2025)
Distraction is All You Need for Multimodal Large Language Model Jailbreaking
von: Yang, Zuopeng, et al.
Veröffentlicht: (2025)
von: Yang, Zuopeng, et al.
Veröffentlicht: (2025)
SO3UFormer: Learning Intrinsic Spherical Features for Rotation-Robust Panoramic Segmentation
von: Zhu, Qinfeng, et al.
Veröffentlicht: (2026)
von: Zhu, Qinfeng, et al.
Veröffentlicht: (2026)
Prosody-Enhanced Acoustic Pre-training and Acoustic-Disentangled Prosody Adapting for Movie Dubbing
von: Zhang, Zhedong, et al.
Veröffentlicht: (2025)
von: Zhang, Zhedong, et al.
Veröffentlicht: (2025)
Enhancing Multimodal Unified Representations for Cross Modal Generalization
von: Huang, Hai, et al.
Veröffentlicht: (2024)
von: Huang, Hai, et al.
Veröffentlicht: (2024)
Open-set Cross Modal Generalization via Multimodal Unified Representation
von: Huang, Hai, et al.
Veröffentlicht: (2025)
von: Huang, Hai, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Causal Disentanglement and Cross-Modal Alignment for Enhanced Few-Shot Learning
von: Jiang, Tianjiao, et al.
Veröffentlicht: (2025) -
Beyond DAGs: A Latent Partial Causal Model for Multimodal Learning
von: Liu, Yuhang, et al.
Veröffentlicht: (2024) -
CLAP: Isolating Content from Style through Contrastive Learning with Augmented Prompts
von: Cai, Yichao, et al.
Veröffentlicht: (2023) -
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
von: Tian, Mingkai, et al.
Veröffentlicht: (2025) -
Concept Component Analysis: A Principled Approach for Concept Extraction in LLMs
von: Liu, Yuhang, et al.
Veröffentlicht: (2026)