Spatial-Semantic Collaborative Cropping for User Generated Content
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Su, Yukun, Cao, Yiwen, Deng, Jingliang, Rao, Fengyun, Wu, Qingyao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
ObjEmbed: Towards Universal Multimodal Object Embeddings
von: Fu, Shenghao, et al.
Veröffentlicht: (2026)
von: Fu, Shenghao, et al.
Veröffentlicht: (2026)
WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization
von: Zhou, Zitang, et al.
Veröffentlicht: (2025)
von: Zhou, Zitang, et al.
Veröffentlicht: (2025)
Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
Unleashing Network Potentials for Semantic Scene Completion
von: Wang, Fengyun, et al.
Veröffentlicht: (2024)
von: Wang, Fengyun, et al.
Veröffentlicht: (2024)
SARA: Controllable Makeup Transfer with Spatial Alignment and Region-Adaptive Normalization
von: Zhong, Xiaojing, et al.
Veröffentlicht: (2023)
von: Zhong, Xiaojing, et al.
Veröffentlicht: (2023)
GPHM: Gaussian Parametric Head Model for Monocular Head Avatar Reconstruction
von: Xu, Yuelang, et al.
Veröffentlicht: (2024)
von: Xu, Yuelang, et al.
Veröffentlicht: (2024)
Content and Salient Semantics Collaboration for Cloth-Changing Person Re-Identification
von: Wang, Qizao, et al.
Veröffentlicht: (2024)
von: Wang, Qizao, et al.
Veröffentlicht: (2024)
D-ORCA: Dialogue-Centric Optimization for Robust Audio-Visual Captioning
von: Tang, Changli, et al.
Veröffentlicht: (2026)
von: Tang, Changli, et al.
Veröffentlicht: (2026)
Semantic-Enriched Latent Visual Reasoning
von: Xu, Tianrun, et al.
Veröffentlicht: (2026)
von: Xu, Tianrun, et al.
Veröffentlicht: (2026)
WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning
von: Yang, Jie, et al.
Veröffentlicht: (2025)
von: Yang, Jie, et al.
Veröffentlicht: (2025)
MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling
von: Yang, Jian, et al.
Veröffentlicht: (2024)
von: Yang, Jian, et al.
Veröffentlicht: (2024)
Revisiting Video Quality Assessment from the Perspective of Generalization
von: Yue, Xinli, et al.
Veröffentlicht: (2024)
von: Yue, Xinli, et al.
Veröffentlicht: (2024)
Video Anomaly Detection with Semantics-Aware Information Bottleneck
von: Li, Juntong, et al.
Veröffentlicht: (2025)
von: Li, Juntong, et al.
Veröffentlicht: (2025)
Multi-Modal Generative Embedding Model
von: Ma, Feipeng, et al.
Veröffentlicht: (2024)
von: Ma, Feipeng, et al.
Veröffentlicht: (2024)
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
von: Zhang, Yunzhu, et al.
Veröffentlicht: (2025)
Instruction-Oriented Preference Alignment for Enhancing Multi-Modal Comprehension Capability of MLLMs
von: Wang, Zitian, et al.
Veröffentlicht: (2025)
von: Wang, Zitian, et al.
Veröffentlicht: (2025)
GenMAC: Compositional Text-to-Video Generation with Multi-Agent Collaboration
von: Huang, Kaiyi, et al.
Veröffentlicht: (2024)
von: Huang, Kaiyi, et al.
Veröffentlicht: (2024)
DepthCropSeg++: Scaling a Crop Segmentation Foundation Model With Depth-Labeled Data
von: Zhang, Jiafei, et al.
Veröffentlicht: (2026)
von: Zhang, Jiafei, et al.
Veröffentlicht: (2026)
R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization
von: Yang, Yi, et al.
Veröffentlicht: (2025)
von: Yang, Yi, et al.
Veröffentlicht: (2025)
MCCD: Multi-Agent Collaboration-based Compositional Diffusion for Complex Text-to-Image Generation
von: Li, Mingcheng, et al.
Veröffentlicht: (2025)
von: Li, Mingcheng, et al.
Veröffentlicht: (2025)
GazeGen: Gaze-Driven User Interaction for Visual Content Generation
von: Hsieh, He-Yen, et al.
Veröffentlicht: (2024)
von: Hsieh, He-Yen, et al.
Veröffentlicht: (2024)
Number it: Temporal Grounding Videos like Flipping Manga
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
von: Wu, Yongliang, et al.
Veröffentlicht: (2024)
An Automated Deep Segmentation and Spatial-Statistics Approach for Post-Blast Rock Fragmentation Assessment
von: Yang, Yukun
Veröffentlicht: (2025)
von: Yang, Yukun
Veröffentlicht: (2025)
DreamComposer++: Empowering Diffusion Models with Multi-View Conditions for 3D Content Generation
von: Yang, Yunhan, et al.
Veröffentlicht: (2025)
von: Yang, Yunhan, et al.
Veröffentlicht: (2025)
Vision-Language Semantic Grounding for Multi-Domain Crop-Weed Segmentation
von: Hossain, Nazia, et al.
Veröffentlicht: (2026)
von: Hossain, Nazia, et al.
Veröffentlicht: (2026)
Instruction-augmented Multimodal Alignment for Image-Text and Element Matching
von: Yue, Xinli, et al.
Veröffentlicht: (2025)
von: Yue, Xinli, et al.
Veröffentlicht: (2025)
Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning
von: Jiang, Huajie, et al.
Veröffentlicht: (2025)
von: Jiang, Huajie, et al.
Veröffentlicht: (2025)
OmniPro: A Comprehensive Benchmark for Omni-Proactive Streaming Video Understanding
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
von: Zhao, Ruixiang, et al.
Veröffentlicht: (2026)
REVERSE: Reinforcing Evidence Verification and Search for Agentic Image geo-localization
von: Li, Yong, et al.
Veröffentlicht: (2026)
von: Li, Yong, et al.
Veröffentlicht: (2026)
HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
von: Wei, Zhixiang, et al.
Veröffentlicht: (2025)
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
von: Tang, Changli, et al.
Veröffentlicht: (2025)
von: Tang, Changli, et al.
Veröffentlicht: (2025)
WeMMU: Enhanced Bridging of Vision-Language Models and Diffusion Models via Noisy Query Tokens
von: Yang, Jian, et al.
Veröffentlicht: (2025)
von: Yang, Jian, et al.
Veröffentlicht: (2025)
Pseudo-Labeling by Multi-Policy Viewfinder Network for Image Cropping
von: Pan, Zhiyu, et al.
Veröffentlicht: (2024)
von: Pan, Zhiyu, et al.
Veröffentlicht: (2024)
SkySeg: Collaborative Onboard Semantic Segmentation with Heterogeneous UAVs in the Wild
von: Lu, Anqi, et al.
Veröffentlicht: (2026)
von: Lu, Anqi, et al.
Veröffentlicht: (2026)
SITSMamba for Crop Classification based on Satellite Image Time Series
von: Qin, Xiaolei, et al.
Veröffentlicht: (2024)
von: Qin, Xiaolei, et al.
Veröffentlicht: (2024)
Semantic and Visual Crop-Guided Diffusion Models for Heterogeneous Tissue Synthesis in Histopathology
von: Alfasly, Saghir, et al.
Veröffentlicht: (2025)
von: Alfasly, Saghir, et al.
Veröffentlicht: (2025)
OmniPart: Part-Aware 3D Generation with Semantic Decoupling and Structural Cohesion
von: Yang, Yunhan, et al.
Veröffentlicht: (2025)
von: Yang, Yunhan, et al.
Veröffentlicht: (2025)
GCP: Guarded Collaborative Perception with Spatial-Temporal Aware Malicious Agent Detection
von: Tao, Yihang, et al.
Veröffentlicht: (2025)
von: Tao, Yihang, et al.
Veröffentlicht: (2025)
SViMo: Synchronized Diffusion for Video and Motion Generation in Hand-object Interaction Scenarios
von: Dang, Lingwei, et al.
Veröffentlicht: (2025)
von: Dang, Lingwei, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
ObjEmbed: Towards Universal Multimodal Object Embeddings
von: Fu, Shenghao, et al.
Veröffentlicht: (2026) -
WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
von: Fu, Shenghao, et al.
Veröffentlicht: (2025) -
HarmonySet: A Comprehensive Dataset for Understanding Video-Music Semantic Alignment and Temporal Synchronization
von: Zhou, Zitang, et al.
Veröffentlicht: (2025) -
Video Repurposing from User Generated Content: A Large-scale Dataset and Benchmark
von: Wu, Yongliang, et al.
Veröffentlicht: (2024) -
Unleashing Network Potentials for Semantic Scene Completion
von: Wang, Fengyun, et al.
Veröffentlicht: (2024)