MMPB: It's Time for Multi-Modal Personalization
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Jaeik, Kim, Woojin, Park, Woohyeon, Do, Jaeyoung |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding
by: Park, Woohyeon, et al.
Published: (2025)
by: Park, Woohyeon, et al.
Published: (2025)
Exploring and Leveraging Class Vectors for Classifier Editing
by: Kim, Jaeik, et al.
Published: (2025)
by: Kim, Jaeik, et al.
Published: (2025)
Vision-Integrated LLMs for Autonomous Driving Assistance : Human Performance Comparison and Trust Evaluation
by: Kim, Namhee, et al.
Published: (2025)
by: Kim, Namhee, et al.
Published: (2025)
MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
by: Park, Woohyeon, et al.
Published: (2026)
by: Park, Woohyeon, et al.
Published: (2026)
PAN-Crafter: Learning Modality-Consistent Alignment for PAN-Sharpening
by: Do, Jeonghyeok, et al.
Published: (2025)
by: Do, Jeonghyeok, et al.
Published: (2025)
Enhanced OoD Detection through Cross-Modal Alignment of Multi-Modal Representations
by: Kim, Jeonghyeon, et al.
Published: (2025)
by: Kim, Jeonghyeon, et al.
Published: (2025)
Real-Time Person Image Synthesis Using a Flow Matching Model
by: Jeong, Jiwoo, et al.
Published: (2025)
by: Jeong, Jiwoo, et al.
Published: (2025)
Generative Phomosaic with Structure-Aligned and Personalized Diffusion
by: Chung, Jaeyoung, et al.
Published: (2026)
by: Chung, Jaeyoung, et al.
Published: (2026)
SEAL-pose: Enhancing 3D Human Pose Estimation via a Learned Loss for Structural Consistency
by: Kim, Yeonsung, et al.
Published: (2026)
by: Kim, Yeonsung, et al.
Published: (2026)
ARGOS: Who, Where, and When in Agentic Multi-Camera Person Search
by: Kim, Myungchul, et al.
Published: (2026)
by: Kim, Myungchul, et al.
Published: (2026)
Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep Allocation
by: Kim, Woojin, et al.
Published: (2025)
by: Kim, Woojin, et al.
Published: (2025)
VizECGNet: Visual ECG Image Network for Cardiovascular Diseases Classification with Multi-Modal Training and Knowledge Distillation
by: Nam, Ju-Hyeon, et al.
Published: (2024)
by: Nam, Ju-Hyeon, et al.
Published: (2024)
FreeTimeGS++: Secrets of Dynamic Gaussian Splatting and Their Principles
by: Lee, Lucas Yunkyu, et al.
Published: (2026)
by: Lee, Lucas Yunkyu, et al.
Published: (2026)
InstructBooth: Instruction-following Personalized Text-to-Image Generation
by: Chae, Daewon, et al.
Published: (2023)
by: Chae, Daewon, et al.
Published: (2023)
ABBSPO: Adaptive Bounding Box Scaling and Symmetric Prior based Orientation Prediction for Detecting Aerial Image Objects
by: Lee, Woojin, et al.
Published: (2025)
by: Lee, Woojin, et al.
Published: (2025)
Missing Modality Prediction for Unpaired Multimodal Learning via Joint Embedding of Unimodal Models
by: Kim, Donggeun, et al.
Published: (2024)
by: Kim, Donggeun, et al.
Published: (2024)
Continual Learning for Multiple Modalities
by: Jin, Hyundong, et al.
Published: (2025)
by: Jin, Hyundong, et al.
Published: (2025)
SToRM: Supervised Token Reduction for Multi-modal LLMs toward efficient end-to-end autonomous driving
by: Kim, Seo Hyun, et al.
Published: (2026)
by: Kim, Seo Hyun, et al.
Published: (2026)
SeMoBridge: Semantic Modality Bridge for Efficient Few-Shot Adaptation of CLIP
by: Timmermann, Christoph, et al.
Published: (2025)
by: Timmermann, Christoph, et al.
Published: (2025)
MMeViT: Multi-Modal ensemble ViT for Post-Stroke Rehabilitation Action Recognition
by: Kim, Ye-eun, et al.
Published: (2025)
by: Kim, Ye-eun, et al.
Published: (2025)
Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification
by: Wang, Xiao, et al.
Published: (2026)
by: Wang, Xiao, et al.
Published: (2026)
VisRef: Visual Refocusing while Thinking Improves Test-Time Scaling in Multi-Modal Large Reasoning Models
by: Ghosal, Soumya Suvra, et al.
Published: (2026)
by: Ghosal, Soumya Suvra, et al.
Published: (2026)
FlipConcept: Tuning-Free Multi-Concept Personalization for Text-to-Image Generation
by: Woo, Young Beom, et al.
Published: (2025)
by: Woo, Young Beom, et al.
Published: (2025)
Preserving Multi-Modal Capabilities of Pre-trained VLMs for Improving Vision-Linguistic Compositionality
by: Oh, Youngtaek, et al.
Published: (2024)
by: Oh, Youngtaek, et al.
Published: (2024)
MM-Diff: High-Fidelity Image Personalization via Multi-Modal Condition Integration
by: Wei, Zhichao, et al.
Published: (2024)
by: Wei, Zhichao, et al.
Published: (2024)
Learning Trimodal Relation for Audio-Visual Question Answering with Missing Modality
by: Park, Kyu Ri, et al.
Published: (2024)
by: Park, Kyu Ri, et al.
Published: (2024)
Progress by Pieces: Test-Time Scaling for Autoregressive Image Generation
by: Park, Joonhyung, et al.
Published: (2025)
by: Park, Joonhyung, et al.
Published: (2025)
Decoupling Stability and Plasticity for Multi-Modal Test-Time Adaptation
by: He, Yongbo, et al.
Published: (2026)
by: He, Yongbo, et al.
Published: (2026)
ReaMIL: Reasoning- and Evidence-Aware Multiple Instance Learning for Whole-Slide Histopathology
by: Jung, Hyun Do, et al.
Published: (2026)
by: Jung, Hyun Do, et al.
Published: (2026)
Mitigating Semantic Collapse in Partially Relevant Video Retrieval
by: Moon, WonJun, et al.
Published: (2025)
by: Moon, WonJun, et al.
Published: (2025)
Fusion Embedding for Pose-Guided Person Image Synthesis with Diffusion Model
by: Lee, Donghwna, et al.
Published: (2024)
by: Lee, Donghwna, et al.
Published: (2024)
MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization
by: Kim, Hyung Kyu, et al.
Published: (2025)
by: Kim, Hyung Kyu, et al.
Published: (2025)
EditSplat: Multi-View Fusion and Attention-Guided Optimization for View-Consistent 3D Scene Editing with 3D Gaussian Splatting
by: Lee, Dong In, et al.
Published: (2024)
by: Lee, Dong In, et al.
Published: (2024)
DynASyn: Multi-Subject Personalization Enabling Dynamic Action Synthesis
by: Choi, Yongjin, et al.
Published: (2025)
by: Choi, Yongjin, et al.
Published: (2025)
Federated Cross-Modal Retrieval with Missing Modalities via Semantic Routing and Adapter Personalization
by: Zhou, Hefeng, et al.
Published: (2026)
by: Zhou, Hefeng, et al.
Published: (2026)
CaddieSet: A Golf Swing Dataset with Human Joint Features and Ball Information
by: Jung, Seunghyeon, et al.
Published: (2025)
by: Jung, Seunghyeon, et al.
Published: (2025)
Safety-Guided Flow (SGF): A Unified Framework for Negative Guidance in Safe Generation
by: Kim, Mingyu, et al.
Published: (2026)
by: Kim, Mingyu, et al.
Published: (2026)
ViTA-PAR: Visual and Textual Attribute Alignment with Attribute Prompting for Pedestrian Attribute Recognition
by: Park, Minjeong, et al.
Published: (2025)
by: Park, Minjeong, et al.
Published: (2025)
Think as Needed: Geometry-Driven Adaptive Perception for Autonomous Driving
by: Kim, Donghyun, et al.
Published: (2026)
by: Kim, Donghyun, et al.
Published: (2026)
VideoRFSplat: Direct Scene-Level Text-to-3D Gaussian Splatting Generation with Flexible Pose and Multi-View Joint Modeling
by: Go, Hyojun, et al.
Published: (2025)
by: Go, Hyojun, et al.
Published: (2025)
Similar Items
-
SECOND: Mitigating Perceptual Hallucination in Vision-Language Models via Selective and Contrastive Decoding
by: Park, Woohyeon, et al.
Published: (2025) -
Exploring and Leveraging Class Vectors for Classifier Editing
by: Kim, Jaeik, et al.
Published: (2025) -
Vision-Integrated LLMs for Autonomous Driving Assistance : Human Performance Comparison and Trust Evaluation
by: Kim, Namhee, et al.
Published: (2025) -
MEDIC-AD: Towards Medical Vision-Language Model's Clinical Intelligence
by: Park, Woohyeon, et al.
Published: (2026) -
PAN-Crafter: Learning Modality-Consistent Alignment for PAN-Sharpening
by: Do, Jeonghyeok, et al.
Published: (2025)