Salvato in:
| Autori principali: | Zhang, Suoxiang, Li, Xiaxi, Chang, Hongrui, Hou, Zhuoyan, Wu, Guoxin, Ji, Ronghua |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2507.05621 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HiGS: Hierarchical Generative Scene Framework for Multi-Step Associative Semantic Spatial Composition
di: Hong, Jiacheng, et al.
Pubblicazione: (2025)
di: Hong, Jiacheng, et al.
Pubblicazione: (2025)
Semantically Consistent Person Image Generation
di: Roy, Prasun, et al.
Pubblicazione: (2023)
di: Roy, Prasun, et al.
Pubblicazione: (2023)
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
di: Wang, Sen, et al.
Pubblicazione: (2025)
di: Wang, Sen, et al.
Pubblicazione: (2025)
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
di: Liu, Ke, et al.
Pubblicazione: (2025)
di: Liu, Ke, et al.
Pubblicazione: (2025)
Visual Semantic Description Generation with MLLMs for Image-Text Matching
di: Chen, Junyu, et al.
Pubblicazione: (2025)
di: Chen, Junyu, et al.
Pubblicazione: (2025)
LipGen: Viseme-Guided Lip Video Generation for Enhancing Visual Speech Recognition
di: Hao, Bowen, et al.
Pubblicazione: (2025)
di: Hao, Bowen, et al.
Pubblicazione: (2025)
CIV-DG: Conditional Instrumental Variables for Domain Generalization in Medical Imaging
di: Bai, Shaojin, et al.
Pubblicazione: (2026)
di: Bai, Shaojin, et al.
Pubblicazione: (2026)
SafePaint: Anti-forensic Image Inpainting with Domain Adaptation
di: Chen, Dunyun, et al.
Pubblicazione: (2024)
di: Chen, Dunyun, et al.
Pubblicazione: (2024)
PersonaGest: Personalized Co-Speech Gesture Generation with Semantic-Guided Hierarchical Motion Representation
di: Zhao, Junchuan, et al.
Pubblicazione: (2026)
di: Zhao, Junchuan, et al.
Pubblicazione: (2026)
Self-distilled Dynamic Fusion Network for Language-based Fashion Retrieval
di: Wu, Yiming, et al.
Pubblicazione: (2024)
di: Wu, Yiming, et al.
Pubblicazione: (2024)
MaVEn: An Effective Multi-granularity Hybrid Visual Encoding Framework for Multimodal Large Language Model
di: Jiang, Chaoya, et al.
Pubblicazione: (2024)
di: Jiang, Chaoya, et al.
Pubblicazione: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
di: Zhang, Zhenxing, et al.
Pubblicazione: (2024)
di: Zhang, Zhenxing, et al.
Pubblicazione: (2024)
Scene Aware Person Image Generation through Global Contextual Conditioning
di: Roy, Prasun, et al.
Pubblicazione: (2022)
di: Roy, Prasun, et al.
Pubblicazione: (2022)
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation
di: Zhang, Xuesong, et al.
Pubblicazione: (2024)
di: Zhang, Xuesong, et al.
Pubblicazione: (2024)
D2SL: Decouple Defogging and Semantic Learning for Foggy Domain-Adaptive Segmentation
di: Sun, Xuan, et al.
Pubblicazione: (2024)
di: Sun, Xuan, et al.
Pubblicazione: (2024)
MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces
di: E, Shaojun, et al.
Pubblicazione: (2025)
di: E, Shaojun, et al.
Pubblicazione: (2025)
Towards Real-World Adverse Weather Image Restoration: Enhancing Clearness and Semantics with Vision-Language Models
di: Xu, Jiaqi, et al.
Pubblicazione: (2024)
di: Xu, Jiaqi, et al.
Pubblicazione: (2024)
Face Consistency Benchmark for GenAI Video
di: Podstawski, Michal, et al.
Pubblicazione: (2025)
di: Podstawski, Michal, et al.
Pubblicazione: (2025)
MHAD: Multimodal Home Activity Dataset with Multi-Angle Videos and Synchronized Physiological Signals
di: Yu, Lei, et al.
Pubblicazione: (2024)
di: Yu, Lei, et al.
Pubblicazione: (2024)
Towards Open-Vocabulary Remote Sensing Image Semantic Segmentation
di: Ye, Chengyang, et al.
Pubblicazione: (2024)
di: Ye, Chengyang, et al.
Pubblicazione: (2024)
MultiSoundGen: Video-to-Audio Generation for Multi-Event Scenarios via SlowFast Contrastive Audio-Visual Pretraining and Direct Preference Optimization
di: Yang, Jianxuan, et al.
Pubblicazione: (2025)
di: Yang, Jianxuan, et al.
Pubblicazione: (2025)
Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval
di: Yang, Yuxin, et al.
Pubblicazione: (2026)
di: Yang, Yuxin, et al.
Pubblicazione: (2026)
G-Refine: A General Quality Refiner for Text-to-Image Generation
di: Li, Chunyi, et al.
Pubblicazione: (2024)
di: Li, Chunyi, et al.
Pubblicazione: (2024)
Multimodal Large Language Model is a Human-Aligned Annotator for Text-to-Image Generation
di: Wu, Xun, et al.
Pubblicazione: (2024)
di: Wu, Xun, et al.
Pubblicazione: (2024)
Exploring Phrase-Level Grounding with Text-to-Image Diffusion Model
di: Yang, Danni, et al.
Pubblicazione: (2024)
di: Yang, Danni, et al.
Pubblicazione: (2024)
Leveraging multimodal explanatory annotations for video interpretation with Modality Specific Dataset
di: Ancarani, Elisa, et al.
Pubblicazione: (2025)
di: Ancarani, Elisa, et al.
Pubblicazione: (2025)
A Unit Enhancement and Guidance Framework for Audio-Driven Avatar Video Generation
di: Zhou, S. Z., et al.
Pubblicazione: (2025)
di: Zhou, S. Z., et al.
Pubblicazione: (2025)
Towards Robust Multimodal Emotion Recognition under Missing Modalities and Distribution Shifts
di: Zhong, Guowei, et al.
Pubblicazione: (2025)
di: Zhong, Guowei, et al.
Pubblicazione: (2025)
Enhancing Environmental Monitoring through Multispectral Imaging: The WasteMS Dataset for Semantic Segmentation of Lakeside Waste
di: Zhu, Qinfeng, et al.
Pubblicazione: (2024)
di: Zhu, Qinfeng, et al.
Pubblicazione: (2024)
MotionPro: A Precise Motion Controller for Image-to-Video Generation
di: Zhang, Zhongwei, et al.
Pubblicazione: (2025)
di: Zhang, Zhongwei, et al.
Pubblicazione: (2025)
DIVE: Inverting Conditional Diffusion Models for Discriminative Tasks
di: Li, Yinqi, et al.
Pubblicazione: (2025)
di: Li, Yinqi, et al.
Pubblicazione: (2025)
Bridging the Pose-Semantic Gap: A Cascade Framework for Text-Based Person Anomaly Search
di: Xie, Zequn, et al.
Pubblicazione: (2026)
di: Xie, Zequn, et al.
Pubblicazione: (2026)
ELIQ: A Label-Free Framework for Quality Assessment of Evolving AI-Generated Images
di: Li, Xinyue, et al.
Pubblicazione: (2026)
di: Li, Xinyue, et al.
Pubblicazione: (2026)
G4G:A Generic Framework for High Fidelity Talking Face Generation with Fine-grained Intra-modal Alignment
di: Zhang, Juan, et al.
Pubblicazione: (2024)
di: Zhang, Juan, et al.
Pubblicazione: (2024)
Hierarchical Textual Knowledge for Enhanced Image Clustering
di: Zhong, Yijie, et al.
Pubblicazione: (2026)
di: Zhong, Yijie, et al.
Pubblicazione: (2026)
Image is All You Need to Empower Large-scale Diffusion Models for In-Domain Generation
di: Cao, Pu, et al.
Pubblicazione: (2023)
di: Cao, Pu, et al.
Pubblicazione: (2023)
MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation
di: Xu, Shuolin, et al.
Pubblicazione: (2025)
di: Xu, Shuolin, et al.
Pubblicazione: (2025)
Optimized Learned Image Compression for Facial Expression Recognition
di: Li, Xiumei, et al.
Pubblicazione: (2025)
di: Li, Xiumei, et al.
Pubblicazione: (2025)
Hierarchical Action Recognition: A Contrastive Video-Language Approach with Hierarchical Interactions
di: Zhang, Rui, et al.
Pubblicazione: (2024)
di: Zhang, Rui, et al.
Pubblicazione: (2024)
KAN-Based Fusion of Dual-Domain for Audio-Driven Facial Landmarks Generation
di: Vo-Thanh, Hoang-Son, et al.
Pubblicazione: (2024)
di: Vo-Thanh, Hoang-Son, et al.
Pubblicazione: (2024)
Documenti analoghi
-
HiGS: Hierarchical Generative Scene Framework for Multi-Step Associative Semantic Spatial Composition
di: Hong, Jiacheng, et al.
Pubblicazione: (2025) -
Semantically Consistent Person Image Generation
di: Roy, Prasun, et al.
Pubblicazione: (2023) -
PanoGen++: Domain-Adapted Text-Guided Panoramic Environment Generation for Vision-and-Language Navigation
di: Wang, Sen, et al.
Pubblicazione: (2025) -
Learning Generalizable and Efficient Image Watermarking via Hierarchical Two-Stage Optimization
di: Liu, Ke, et al.
Pubblicazione: (2025) -
Visual Semantic Description Generation with MLLMs for Image-Text Matching
di: Chen, Junyu, et al.
Pubblicazione: (2025)