Gespeichert in:
| Hauptverfasser: | Liu, Hongyuan, Yang, Qinli, Li, Wen, Zhang, Zhong, Liu, Jiaming, Han, Wei, Qin, Zhili, Guo, Jinxia, Shao, Junming |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2604.00279 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Exploiting Fine-Grained Prototype Distribution for Boosting Unsupervised Class Incremental Learning
von: Liu, Jiaming, et al.
Veröffentlicht: (2024)
von: Liu, Jiaming, et al.
Veröffentlicht: (2024)
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025)
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025)
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
von: Wang, Zitian, et al.
Veröffentlicht: (2025)
von: Wang, Zitian, et al.
Veröffentlicht: (2025)
Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-Guidance
von: Hur, Jiwan, et al.
Veröffentlicht: (2024)
von: Hur, Jiwan, et al.
Veröffentlicht: (2024)
Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2026)
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2026)
Geo-Align: Video Generation Alignment via Metric Geometry Reward
von: Li, Zizun, et al.
Veröffentlicht: (2026)
von: Li, Zizun, et al.
Veröffentlicht: (2026)
AlignGen: Boosting Personalized Image Generation with Cross-Modality Prior Alignment
von: Lin, Yiheng, et al.
Veröffentlicht: (2025)
von: Lin, Yiheng, et al.
Veröffentlicht: (2025)
ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
von: Yang, Cheng, et al.
Veröffentlicht: (2024)
Enhancing Compositional Generalization via Compositional Feature Alignment
von: Wang, Haoxiang, et al.
Veröffentlicht: (2024)
von: Wang, Haoxiang, et al.
Veröffentlicht: (2024)
Geometry-Aware Cross Modal Alignment for Light Field-LiDAR Semantic Segmentation
von: Luo, Jie, et al.
Veröffentlicht: (2025)
von: Luo, Jie, et al.
Veröffentlicht: (2025)
Integrating MedCLIP and Cross-Modal Fusion for Automatic Radiology Report Generation
von: Han, Qianhao, et al.
Veröffentlicht: (2024)
von: Han, Qianhao, et al.
Veröffentlicht: (2024)
OmniJigsaw: Enhancing Omni-Modal Reasoning via Modality-Orchestrated Reordering
von: Jia, Yiduo, et al.
Veröffentlicht: (2026)
von: Jia, Yiduo, et al.
Veröffentlicht: (2026)
Beyond Loss Values: Robust Dynamic Pruning via Loss Trajectory Alignment
von: Qin, Huaiyuan, et al.
Veröffentlicht: (2026)
von: Qin, Huaiyuan, et al.
Veröffentlicht: (2026)
DM$^3$T: Harmonizing Modalities via Diffusion for Multi-Object Tracking
von: Li, Weiran, et al.
Veröffentlicht: (2025)
von: Li, Weiran, et al.
Veröffentlicht: (2025)
Optimizing Distributional Geometry Alignment with Optimal Transport for Generative Dataset Distillation
von: Cui, Xiao, et al.
Veröffentlicht: (2025)
von: Cui, Xiao, et al.
Veröffentlicht: (2025)
VIGOR: VIdeo Geometry-Oriented Reward for Temporal Generative Alignment
von: Yin, Tengjiao, et al.
Veröffentlicht: (2026)
von: Yin, Tengjiao, et al.
Veröffentlicht: (2026)
Unlocking Multilingual Reasoning Capability of LLMs and LVLMs through Representation Engineering
von: Li, Qiming, et al.
Veröffentlicht: (2025)
von: Li, Qiming, et al.
Veröffentlicht: (2025)
M2I2HA: Multi-modal Object Detection Based on Intra- and Inter-Modal Hypergraph Attention
von: Yang, Xiaofan, et al.
Veröffentlicht: (2026)
von: Yang, Xiaofan, et al.
Veröffentlicht: (2026)
UniFork: Exploring Modality Alignment for Unified Multimodal Understanding and Generation
von: Li, Teng, et al.
Veröffentlicht: (2025)
von: Li, Teng, et al.
Veröffentlicht: (2025)
Quantum-Inspired Spectral Geometry for Neural Operator Equivalence and Structured Pruning
von: Shao, Haijian, et al.
Veröffentlicht: (2025)
von: Shao, Haijian, et al.
Veröffentlicht: (2025)
Unlocking the Latent Canvas: Eliciting and Benchmarking Symbolic Visual Expression in LLMs
von: Zheng, Yiren, et al.
Veröffentlicht: (2026)
von: Zheng, Yiren, et al.
Veröffentlicht: (2026)
Symbiotic-MoE: Unlocking the Synergy between Generation and Understanding
von: Liu, Xiangyue, et al.
Veröffentlicht: (2026)
von: Liu, Xiangyue, et al.
Veröffentlicht: (2026)
GIDE: Unlocking Diffusion LLMs for Precise Training-Free Image Editing
von: Zhu, Zifeng, et al.
Veröffentlicht: (2026)
von: Zhu, Zifeng, et al.
Veröffentlicht: (2026)
Enhance Vision-Language Alignment with Noise
von: Huang, Sida, et al.
Veröffentlicht: (2024)
von: Huang, Sida, et al.
Veröffentlicht: (2024)
Robust Tracking via Mamba-based Context-aware Token Learning
von: Xie, Jinxia, et al.
Veröffentlicht: (2024)
von: Xie, Jinxia, et al.
Veröffentlicht: (2024)
Cross Modal Fine-Grained Alignment via Granularity-Aware and Region-Uncertain Modeling
von: Liu, Jiale, et al.
Veröffentlicht: (2025)
von: Liu, Jiale, et al.
Veröffentlicht: (2025)
Cross-Modal Scene Semantic Alignment for Image Complexity Assessment
von: Luo, Yuqing, et al.
Veröffentlicht: (2025)
von: Luo, Yuqing, et al.
Veröffentlicht: (2025)
Attribute-Aware Implicit Modality Alignment for Text Attribute Person Search
von: Wang, Xin, et al.
Veröffentlicht: (2024)
von: Wang, Xin, et al.
Veröffentlicht: (2024)
General-Purpose Multi-Modal OOD Detection Framework
von: Duong, Viet, et al.
Veröffentlicht: (2023)
von: Duong, Viet, et al.
Veröffentlicht: (2023)
Investigating and Enhancing Vision-Audio Capability in Omnimodal Large Language Models
von: Hu, Rui, et al.
Veröffentlicht: (2025)
von: Hu, Rui, et al.
Veröffentlicht: (2025)
MMGDreamer: Mixed-Modality Graph for Geometry-Controllable 3D Indoor Scene Generation
von: Yang, Zhifei, et al.
Veröffentlicht: (2025)
von: Yang, Zhifei, et al.
Veröffentlicht: (2025)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
von: Zhang, Zhenxing, et al.
Veröffentlicht: (2024)
SpatialFly: Geometry-Guided Representation Alignment for UAV Vision-and-Language Navigation in Urban Environments
von: Jiang, Wen, et al.
Veröffentlicht: (2026)
von: Jiang, Wen, et al.
Veröffentlicht: (2026)
Unlocking Few-Shot Capabilities in LVLMs via Prompt Conditioning and Head Selection
von: de Senneville, Adhemar, et al.
Veröffentlicht: (2026)
von: de Senneville, Adhemar, et al.
Veröffentlicht: (2026)
MegActor-$Σ$: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer
von: Yang, Shurong, et al.
Veröffentlicht: (2024)
von: Yang, Shurong, et al.
Veröffentlicht: (2024)
MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
von: Wang, Tianyi, et al.
Veröffentlicht: (2025)
Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment
von: Zhao, Pengfei, et al.
Veröffentlicht: (2025)
von: Zhao, Pengfei, et al.
Veröffentlicht: (2025)
GeoWorld: Unlocking the Potential of Geometry Models to Facilitate High-Fidelity 3D Scene Generation
von: Wan, Yuhao, et al.
Veröffentlicht: (2025)
von: Wan, Yuhao, et al.
Veröffentlicht: (2025)
PRISM2: Unlocking Multi-Modal General Pathology AI with Clinical Dialogue
von: Vorontsov, Eugene, et al.
Veröffentlicht: (2025)
von: Vorontsov, Eugene, et al.
Veröffentlicht: (2025)
SafeSora: Towards Safety Alignment of Text2Video Generation via a Human Preference Dataset
von: Dai, Josef, et al.
Veröffentlicht: (2024)
von: Dai, Josef, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Exploiting Fine-Grained Prototype Distribution for Boosting Unsupervised Class Incremental Learning
von: Liu, Jiaming, et al.
Veröffentlicht: (2024) -
3DWG: 3D Weakly Supervised Visual Grounding via Category and Instance-Level Alignment
von: Li, Xiaoqi, et al.
Veröffentlicht: (2025) -
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
von: Wang, Zitian, et al.
Veröffentlicht: (2025) -
Unlocking the Capabilities of Masked Generative Models for Image Synthesis via Self-Guidance
von: Hur, Jiwan, et al.
Veröffentlicht: (2024) -
Geometry-Aware CLIP Retrieval via Local Cross-Modal Alignment and Steering
von: Prakash, Nirmalendu, et al.
Veröffentlicht: (2026)