Gespeichert in:
| Hauptverfasser: | Zheng, Jiawen, Jia, Haonan, Li, Ming, Zheng, Yuhui, Zeng, Yufeng, Gao, Yang, Liang, Chen |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2602.07495 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HMPE:HeatMap Embedding for Efficient Transformer-Based Small Object Detection
von: Zeng, YangChen
Veröffentlicht: (2025)
von: Zeng, YangChen
Veröffentlicht: (2025)
Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026)
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation
von: Zhang, Xuesong, et al.
Veröffentlicht: (2024)
von: Zhang, Xuesong, et al.
Veröffentlicht: (2024)
Noise-Tolerant Learning for Audio-Visual Action Recognition
von: Han, Haochen, et al.
Veröffentlicht: (2022)
von: Han, Haochen, et al.
Veröffentlicht: (2022)
GSSF: Generalized Structural Sparse Function for Deep Cross-modal Metric Learning
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)
Zero-Shot Visual Grounding in 3D Gaussians via View Retrieval
von: Liao, Liwei, et al.
Veröffentlicht: (2025)
von: Liao, Liwei, et al.
Veröffentlicht: (2025)
Wills Aligner: Multi-Subject Collaborative Brain Visual Decoding
von: Bao, Guangyin, et al.
Veröffentlicht: (2024)
von: Bao, Guangyin, et al.
Veröffentlicht: (2024)
PathVLM-R1: A Reinforcement Learning-Driven Reasoning Model for Pathology Visual-Language Tasks
von: Wu, Jianyu, et al.
Veröffentlicht: (2025)
von: Wu, Jianyu, et al.
Veröffentlicht: (2025)
ROI-Guided Point Cloud Geometry Compression Towards Human and Machine Vision
von: Liang, Xie, et al.
Veröffentlicht: (2025)
von: Liang, Xie, et al.
Veröffentlicht: (2025)
Hierarchical Masked Autoregressive Models with Low-Resolution Token Pivots
von: Zheng, Guangting, et al.
Veröffentlicht: (2025)
von: Zheng, Guangting, et al.
Veröffentlicht: (2025)
Hierarchical Sub-action Tree for Continuous Sign Language Recognition
von: Yang, Dejie, et al.
Veröffentlicht: (2025)
von: Yang, Dejie, et al.
Veröffentlicht: (2025)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
von: Chen, Junyu, et al.
Veröffentlicht: (2025)
MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
von: Zheng, Junjie, et al.
Veröffentlicht: (2025)
PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
von: Zhao, Junchuan, et al.
Veröffentlicht: (2026)
von: Zhao, Junchuan, et al.
Veröffentlicht: (2026)
Towards Universal Modal Tracking with Online Dense Temporal Token Learning
von: Zheng, Yaozong, et al.
Veröffentlicht: (2025)
von: Zheng, Yaozong, et al.
Veröffentlicht: (2025)
HCNQA: Enhancing 3D VQA with Hierarchical Concentration Narrowing Supervision
von: Zhou, Shengli, et al.
Veröffentlicht: (2025)
von: Zhou, Shengli, et al.
Veröffentlicht: (2025)
Towards Flexible Evaluation for Generative Visual Question Answering
von: Ji, Huishan, et al.
Veröffentlicht: (2024)
von: Ji, Huishan, et al.
Veröffentlicht: (2024)
OS-HGAdapter: Open Semantic Hypergraph Adapter for Large Language Models Assisted Entropy-Enhanced Image-Text Alignment
von: Chen, Rongjun, et al.
Veröffentlicht: (2025)
von: Chen, Rongjun, et al.
Veröffentlicht: (2025)
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
von: Liu, Ke, et al.
Veröffentlicht: (2025)
von: Liu, Ke, et al.
Veröffentlicht: (2025)
Learning Video Context as Interleaved Multimodal Sequences
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2024)
SonoWorld: From One Image to a 3D Audio-Visual Scene
von: Jin, Derong, et al.
Veröffentlicht: (2026)
von: Jin, Derong, et al.
Veröffentlicht: (2026)
Teacher-Guided Pseudo Supervision and Cross-Modal Alignment for Audio-Visual Video Parsing
von: Chen, Yaru, et al.
Veröffentlicht: (2025)
von: Chen, Yaru, et al.
Veröffentlicht: (2025)
DIRECT: Video Mashup Creation via Hierarchical Multi-Agent Planning and Intent-Guided Editing
von: Li, Ke, et al.
Veröffentlicht: (2026)
von: Li, Ke, et al.
Veröffentlicht: (2026)
Delving Deeper: Hierarchical Visual Perception for Robust Video-Text Retrieval
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
von: Xie, Zequn, et al.
Veröffentlicht: (2026)
Deep Reversible Consistency Learning for Cross-modal Retrieval
von: Pu, Ruitao, et al.
Veröffentlicht: (2025)
von: Pu, Ruitao, et al.
Veröffentlicht: (2025)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
von: Li, Huilai, et al.
Veröffentlicht: (2026)
von: Li, Huilai, et al.
Veröffentlicht: (2026)
Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives
von: Chen, Yi-Chun
Veröffentlicht: (2025)
von: Chen, Yi-Chun
Veröffentlicht: (2025)
Beyond Walking: A Large-Scale Image-Text Benchmark for Text-based Person Anomaly Search
von: Yang, Shuyu, et al.
Veröffentlicht: (2024)
von: Yang, Shuyu, et al.
Veröffentlicht: (2024)
Attributes-aware Visual Emotion Representation Learning
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025)
von: Maharjan, Rahul Singh, et al.
Veröffentlicht: (2025)
Factorized Learning for Temporally Grounded Video-Language Models
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
von: Zeng, Wenzheng, et al.
Veröffentlicht: (2025)
ASR-enhanced Multimodal Representation Learning for Cross-Domain Product Retrieval
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2024)
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2024)
Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation
von: Cai, Haonan, et al.
Veröffentlicht: (2026)
von: Cai, Haonan, et al.
Veröffentlicht: (2026)
Hypergraph Tversky-Aware Domain Incremental Learning for Brain Tumor Segmentation with Missing Modalities
von: Wang, Junze, et al.
Veröffentlicht: (2025)
von: Wang, Junze, et al.
Veröffentlicht: (2025)
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
von: Lu, Xingyu, et al.
Veröffentlicht: (2026)
von: Lu, Xingyu, et al.
Veröffentlicht: (2026)
Causal Debiasing for Visual Commonsense Reasoning
von: Zou, Jiayi, et al.
Veröffentlicht: (2025)
von: Zou, Jiayi, et al.
Veröffentlicht: (2025)
CAMeL: Cross-modality Adaptive Meta-Learning for Text-based Person Retrieval
von: Yu, Hang, et al.
Veröffentlicht: (2025)
von: Yu, Hang, et al.
Veröffentlicht: (2025)
Loc4Plan: Locating Before Planning for Outdoor Vision and Language Navigation
von: Tian, Huilin, et al.
Veröffentlicht: (2024)
von: Tian, Huilin, et al.
Veröffentlicht: (2024)
Language-Guided Diffusion Model for Visual Grounding
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
von: Chen, Sijia, et al.
Veröffentlicht: (2023)
LinVT: Empower Your Image-level Large Language Model to Understand Videos
von: Gao, Lishuai, et al.
Veröffentlicht: (2024)
von: Gao, Lishuai, et al.
Veröffentlicht: (2024)
CustomContrast: A Multilevel Contrastive Perspective For Subject-Driven Text-to-Image Customization
von: Chen, Nan, et al.
Veröffentlicht: (2024)
von: Chen, Nan, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
HMPE:HeatMap Embedding for Efficient Transformer-Based Small Object Detection
von: Zeng, YangChen
Veröffentlicht: (2025) -
Hierarchical Semantic Correlation-Aware Masked Autoencoder for Unsupervised Audio-Visual Representation Learning
von: Zeng, Donghuo, et al.
Veröffentlicht: (2026) -
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation
von: Zhang, Xuesong, et al.
Veröffentlicht: (2024) -
Noise-Tolerant Learning for Audio-Visual Action Recognition
von: Han, Haochen, et al.
Veröffentlicht: (2022) -
GSSF: Generalized Structural Sparse Function for Deep Cross-modal Metric Learning
von: Diao, Haiwen, et al.
Veröffentlicht: (2024)