Saved in:
| Main Authors: | Xia, Weihao, Zhou, Chenliang, Oztireli, Cengiz |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2505.14319 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
FreNBRDF: A Frequency-Rectified Neural Material Representation
by: Zhou, Chenliang, et al.
Published: (2025)
by: Zhou, Chenliang, et al.
Published: (2025)
Exploring The Visual Feature Space for Multimodal Neural Decoding
by: Xia, Weihao, et al.
Published: (2025)
by: Xia, Weihao, et al.
Published: (2025)
Quartet of Diffusions: Structure-Aware Point Cloud Generation through Part and Symmetry Guidance
by: Zhou, Chenliang, et al.
Published: (2026)
by: Zhou, Chenliang, et al.
Published: (2026)
Multigranular Evaluation for Brain Visual Decoding
by: Xia, Weihao, et al.
Published: (2025)
by: Xia, Weihao, et al.
Published: (2025)
PriOr-Flow: Enhancing Primitive Panoramic Optical Flow with Orthogonal View
by: Liu, Longliang, et al.
Published: (2025)
by: Liu, Longliang, et al.
Published: (2025)
CLIP-PAE: Projection-Augmentation Embedding to Extract Relevant Features for a Disentangled, Interpretable, and Controllable Text-Guided Face Manipulation
by: Zhou, Chenliang, et al.
Published: (2022)
by: Zhou, Chenliang, et al.
Published: (2022)
Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models
by: Ye, Weihao, et al.
Published: (2024)
by: Ye, Weihao, et al.
Published: (2024)
Principled Multimodal Representation Learning
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Learning Brain Representation with Hierarchical Visual Embeddings
by: Zheng, Jiawen, et al.
Published: (2026)
by: Zheng, Jiawen, et al.
Published: (2026)
Self-supervised Photographic Image Layout Representation Learning
by: Zhao, Zhaoran, et al.
Published: (2024)
by: Zhao, Zhaoran, et al.
Published: (2024)
Calibrated Multimodal Representation Learning with Missing Modalities
by: Liu, Xiaohao, et al.
Published: (2025)
by: Liu, Xiaohao, et al.
Published: (2025)
Rethinking Multi-view Representation Learning via Distilled Disentangling
by: Ke, Guanzhou, et al.
Published: (2024)
by: Ke, Guanzhou, et al.
Published: (2024)
MSLIQA: Enhancing Learning Representations for Image Quality Assessment through Multi-Scale Learning
by: Avanaki, Nasim Jamshidi, et al.
Published: (2024)
by: Avanaki, Nasim Jamshidi, et al.
Published: (2024)
UniTalking: A Unified Audio-Video Framework for Talking Portrait Generation
by: Li, Hebeizi, et al.
Published: (2026)
by: Li, Hebeizi, et al.
Published: (2026)
Pedestrian Trajectory Prediction Based on Social Interactions Learning With Random Weights
by: Xie, Jiajia, et al.
Published: (2025)
by: Xie, Jiajia, et al.
Published: (2025)
Efficient Object-centric Representation Learning with Pre-trained Geometric Prior
by: Khac, Phúc H. Le, et al.
Published: (2024)
by: Khac, Phúc H. Le, et al.
Published: (2024)
Towards Unbiased Cross-Modal Representation Learning for Food Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
DS-NeRV: Implicit Neural Video Representation with Decomposed Static and Dynamic Codes
by: Yan, Hao, et al.
Published: (2024)
by: Yan, Hao, et al.
Published: (2024)
UAU-Net: Uncertainty-aware Representation Learning and Evidential Classification for Facial Action Unit Detection
by: Li, Yuze, et al.
Published: (2026)
by: Li, Yuze, et al.
Published: (2026)
Bi-VLDoc: Bidirectional Vision-Language Modeling for Visually-Rich Document Understanding
by: Luo, Chuwei, et al.
Published: (2022)
by: Luo, Chuwei, et al.
Published: (2022)
Learning Long-Range Action Representation by Two-Stream Mamba Pyramid Network for Figure Skating Assessment
by: Wang, Fengshun, et al.
Published: (2025)
by: Wang, Fengshun, et al.
Published: (2025)
MIRROR: Multi-Modal Pathological Self-Supervised Representation Learning via Modality Alignment and Retention
by: Wang, Tianyi, et al.
Published: (2025)
by: Wang, Tianyi, et al.
Published: (2025)
Learning Video Context as Interleaved Multimodal Sequences
by: Lin, Kevin Qinghong, et al.
Published: (2024)
by: Lin, Kevin Qinghong, et al.
Published: (2024)
Rethinking Video with a Universal Event-Based Representation
by: Freeman, Andrew
Published: (2024)
by: Freeman, Andrew
Published: (2024)
LocoMotion: Learning Motion-Focused Video-Language Representations
by: Doughty, Hazel, et al.
Published: (2024)
by: Doughty, Hazel, et al.
Published: (2024)
ASAP: Advancing Semantic Alignment Promotes Multi-Modal Manipulation Detecting and Grounding
by: Zhang, Zhenxing, et al.
Published: (2024)
by: Zhang, Zhenxing, et al.
Published: (2024)
Mitigating Cross-modal Representation Bias for Multicultural Image-to-Recipe Retrieval
by: Wang, Qing, et al.
Published: (2025)
by: Wang, Qing, et al.
Published: (2025)
DIVA: Harnessing the Representation Divergence in Unified Multimodal Models for Mutual Reinforcement
by: Lu, Renjie, et al.
Published: (2026)
by: Lu, Renjie, et al.
Published: (2026)
NeR-SC: Adapting Neural Video Representation to Screen Content
by: Shi, Ruohan, et al.
Published: (2026)
by: Shi, Ruohan, et al.
Published: (2026)
GSVC: Efficient Video Representation and Compression Through 2D Gaussian Splatting
by: Wang, Longan, et al.
Published: (2025)
by: Wang, Longan, et al.
Published: (2025)
Sketch and Patch: Efficient 3D Gaussian Representation for Man-Made Scenes
by: Shi, Yuang, et al.
Published: (2025)
by: Shi, Yuang, et al.
Published: (2025)
UMBRAE: Unified Multimodal Brain Decoding
by: Xia, Weihao, et al.
Published: (2024)
by: Xia, Weihao, et al.
Published: (2024)
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing
by: Li, Huilai, et al.
Published: (2026)
by: Li, Huilai, et al.
Published: (2026)
Attributes-aware Visual Emotion Representation Learning
by: Maharjan, Rahul Singh, et al.
Published: (2025)
by: Maharjan, Rahul Singh, et al.
Published: (2025)
Noisy-Correspondence Learning for Text-to-Image Person Re-identification
by: Qin, Yang, et al.
Published: (2023)
by: Qin, Yang, et al.
Published: (2023)
MultiColor: Image Colorization by Learning from Multiple Color Spaces
by: Du, Xiangcheng, et al.
Published: (2024)
by: Du, Xiangcheng, et al.
Published: (2024)
Logic Unseen: Revealing the Logical Blindspots of Vision-Language Models
by: Zhou, Yuchen, et al.
Published: (2025)
by: Zhou, Yuchen, et al.
Published: (2025)
A Unified Non-Parametric and Interpretable Point Cloud Analysis via t-FCW Graph Representation
by: Lai, Haijian, et al.
Published: (2026)
by: Lai, Haijian, et al.
Published: (2026)
Agent Journey Beyond RGB: Hierarchical Semantic-Spatial Representation Enrichment for Vision-and-Language Navigation
by: Zhang, Xuesong, et al.
Published: (2024)
by: Zhang, Xuesong, et al.
Published: (2024)
Learning Spatial Adaptation and Temporal Coherence in Diffusion Models for Video Super-Resolution
by: Chen, Zhikai, et al.
Published: (2024)
by: Chen, Zhikai, et al.
Published: (2024)
Similar Items
-
FreNBRDF: A Frequency-Rectified Neural Material Representation
by: Zhou, Chenliang, et al.
Published: (2025) -
Exploring The Visual Feature Space for Multimodal Neural Decoding
by: Xia, Weihao, et al.
Published: (2025) -
Quartet of Diffusions: Structure-Aware Point Cloud Generation through Part and Symmetry Guidance
by: Zhou, Chenliang, et al.
Published: (2026) -
Multigranular Evaluation for Brain Visual Decoding
by: Xia, Weihao, et al.
Published: (2025) -
PriOr-Flow: Enhancing Primitive Panoramic Optical Flow with Orthogonal View
by: Liu, Longliang, et al.
Published: (2025)