Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP Supervision
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Yin, Kangsheng, Liu, Quan, Shen, Xuelin, He, Yulin, Yang, Wenhan, Wang, Shiqi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Sliced Maximal Information Coefficient: A Training-Free Approach for Image Quality Assessment Enhancement
von: Xiao, Kang, et al.
Veröffentlicht: (2024)
von: Xiao, Kang, et al.
Veröffentlicht: (2024)
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
von: Zhang, Pingping, et al.
Veröffentlicht: (2024)
Privacy-Shielded Image Compression: Defending Against Exploitation from Vision-Language Pretrained Models
von: Shen, Xuelin, et al.
Veröffentlicht: (2025)
von: Shen, Xuelin, et al.
Veröffentlicht: (2025)
FakeBench: Probing Explainable Fake Image Detection via Large Multimodal Models
von: Li, Yixuan, et al.
Veröffentlicht: (2024)
von: Li, Yixuan, et al.
Veröffentlicht: (2024)
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)
Scalable Image Coding for Humans and Machines Using Feature Fusion Network
von: Shindo, Takahiro, et al.
Veröffentlicht: (2024)
von: Shindo, Takahiro, et al.
Veröffentlicht: (2024)
OneHOI: Unifying Human-Object Interaction Generation and Editing
von: Hoe, Jiun Tian, et al.
Veröffentlicht: (2026)
von: Hoe, Jiun Tian, et al.
Veröffentlicht: (2026)
Test-Time Adaptation with CLIP Reward for Zero-Shot Generalization in Vision-Language Models
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
von: Zhao, Shuai, et al.
Veröffentlicht: (2023)
Graph-Driven Multimodal Feature Learning Framework for Apparent Personality Assessment
von: Wang, Kangsheng, et al.
Veröffentlicht: (2025)
von: Wang, Kangsheng, et al.
Veröffentlicht: (2025)
Opinion-Unaware Blind Image Quality Assessment using Multi-Scale Deep Feature Statistics
von: Ni, Zhangkai, et al.
Veröffentlicht: (2024)
von: Ni, Zhangkai, et al.
Veröffentlicht: (2024)
Not Just What's There: Enabling CLIP to Comprehend Negated Visual Descriptions Without Fine-tuning
von: Xiao, Junhao, et al.
Veröffentlicht: (2026)
von: Xiao, Junhao, et al.
Veröffentlicht: (2026)
Unlearning the Noisy Correspondence Makes CLIP More Robust
von: Han, Haochen, et al.
Veröffentlicht: (2025)
von: Han, Haochen, et al.
Veröffentlicht: (2025)
CLIP Brings Better Features to Visual Aesthetics Learners
von: Xu, Liwu, et al.
Veröffentlicht: (2023)
von: Xu, Liwu, et al.
Veröffentlicht: (2023)
Discriminative-Generative Synergy for Occlusion Robust 3D Human Mesh Recovery
von: Liu, Yang, et al.
Veröffentlicht: (2026)
von: Liu, Yang, et al.
Veröffentlicht: (2026)
Pre-training CLIP against Data Poisoning with Optimal Transport-based Matching and Alignment
von: Zhang, Tong, et al.
Veröffentlicht: (2025)
von: Zhang, Tong, et al.
Veröffentlicht: (2025)
HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
von: Chen, Liyang, et al.
Veröffentlicht: (2025)
Multiscale Feature Importance-based Bit Allocation for End-to-End Feature Coding for Machines
von: Liu, Junle, et al.
Veröffentlicht: (2025)
von: Liu, Junle, et al.
Veröffentlicht: (2025)
LMM-Regularized CLIP Embeddings for Image Classification
von: Tzelepi, Maria, et al.
Veröffentlicht: (2024)
von: Tzelepi, Maria, et al.
Veröffentlicht: (2024)
Context-Enhanced Video Moment Retrieval with Large Language Models
von: Liu, Weijia, et al.
Veröffentlicht: (2024)
von: Liu, Weijia, et al.
Veröffentlicht: (2024)
SOSControl: Enhancing Human Motion Generation through Saliency-Aware Symbolic Orientation and Timing Control
von: Au, Ho Yin, et al.
Veröffentlicht: (2025)
von: Au, Ho Yin, et al.
Veröffentlicht: (2025)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
von: Li, Hebeizi, et al.
Veröffentlicht: (2026)
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer
von: Cai, Qi, et al.
Veröffentlicht: (2026)
von: Cai, Qi, et al.
Veröffentlicht: (2026)
Human Motion Video Generation: A Survey
von: Xue, Haiwei, et al.
Veröffentlicht: (2025)
von: Xue, Haiwei, et al.
Veröffentlicht: (2025)
CLIP-PCQA: Exploring Subjective-Aligned Vision-Language Modeling for Point Cloud Quality Assessment
von: Liu, Yating, et al.
Veröffentlicht: (2025)
von: Liu, Yating, et al.
Veröffentlicht: (2025)
Predicting Satisfied User and Machine Ratio for Compressed Images: A Unified Approach
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
von: Zhang, Qi, et al.
Veröffentlicht: (2024)
Selective Vision-Language Subspace Projection for Few-shot CLIP
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
von: Zhu, Xingyu, et al.
Veröffentlicht: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
von: Chow, Wei, et al.
Veröffentlicht: (2024)
von: Chow, Wei, et al.
Veröffentlicht: (2024)
PAME: Self-Supervised Masked Autoencoder for No-Reference Point Cloud Quality Assessment
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
von: Shan, Ziyu, et al.
Veröffentlicht: (2024)
Synthetic Perception: Can Generated Images Unlock Latent Visual Prior for Text-Centric Reasoning?
von: Huang, Yuesheng, et al.
Veröffentlicht: (2025)
von: Huang, Yuesheng, et al.
Veröffentlicht: (2025)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
von: Li, Huilai, et al.
Veröffentlicht: (2026)
von: Li, Huilai, et al.
Veröffentlicht: (2026)
ROI-Guided Point Cloud Geometry Compression Towards Human and Machine Vision
von: Liang, Xie, et al.
Veröffentlicht: (2025)
von: Liang, Xie, et al.
Veröffentlicht: (2025)
ReSyncer: Rewiring Style-based Generator for Unified Audio-Visually Synced Facial Performer
von: Guan, Jiazhi, et al.
Veröffentlicht: (2024)
von: Guan, Jiazhi, et al.
Veröffentlicht: (2024)
M3FAS: An Accurate and Robust MultiModal Mobile Face Anti-Spoofing System
von: Kong, Chenqi, et al.
Veröffentlicht: (2023)
von: Kong, Chenqi, et al.
Veröffentlicht: (2023)
EDGE-Shield: Efficient Denoising-staGE Shield for Violative Content Filtering via Scalable Reference-Based Matching
von: Taniguchi, Takara, et al.
Veröffentlicht: (2026)
von: Taniguchi, Takara, et al.
Veröffentlicht: (2026)
Bridging Your Imagination with Audio-Video Generation via a Unified Director
von: Zhang, Jiaxu, et al.
Veröffentlicht: (2025)
von: Zhang, Jiaxu, et al.
Veröffentlicht: (2025)
CUS3D :CLIP-based Unsupervised 3D Segmentation via Object-level Denoise
von: Yu, Fuyang, et al.
Veröffentlicht: (2024)
von: Yu, Fuyang, et al.
Veröffentlicht: (2024)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
von: Wang, Jiapeng, et al.
Veröffentlicht: (2024)
JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
von: Liu, Kai, et al.
Veröffentlicht: (2026)
von: Liu, Kai, et al.
Veröffentlicht: (2026)
T-GVC: Trajectory-Guided Generative Video Coding at Ultra-Low Bitrates
von: Wang, Zhitao, et al.
Veröffentlicht: (2025)
von: Wang, Zhitao, et al.
Veröffentlicht: (2025)
InstructHumans: Editing Animated 3D Human Textures with Instructions
von: Zhu, Jiayin, et al.
Veröffentlicht: (2024)
von: Zhu, Jiayin, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Sliced Maximal Information Coefficient: A Training-Free Approach for Image Quality Assessment Enhancement
von: Xiao, Kang, et al.
Veröffentlicht: (2024) -
When Video Coding Meets Multimodal Large Language Models: A Unified Paradigm for Video Coding
von: Zhang, Pingping, et al.
Veröffentlicht: (2024) -
Privacy-Shielded Image Compression: Defending Against Exploitation from Vision-Language Pretrained Models
von: Shen, Xuelin, et al.
Veröffentlicht: (2025) -
FakeBench: Probing Explainable Fake Image Detection via Large Multimodal Models
von: Li, Yixuan, et al.
Veröffentlicht: (2024) -
UniCode$^2$: Cascaded Large-scale Codebooks for Unified Multimodal Understanding and Generation
von: Chen, Yanzhe, et al.
Veröffentlicht: (2025)