Omni-Persona: Systematic Benchmarking and Improving Omnimodal Personalization
Fuente:
arXiv
Saved in:
| Main Authors: | Oh, Yeongtak, Lee, Dongwook, Park, Sangkwon, Kim, Heeseung, Yoon, Sungroh |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Style-Friendly SNR Sampler for Style-Driven Generation
by: Choi, Jooyoung, et al.
Published: (2024)
by: Choi, Jooyoung, et al.
Published: (2024)
On mitigating stability-plasticity dilemma in CLIP-guided image morphing via geodesic distillation loss
by: Oh, Yeongtak, et al.
Published: (2024)
by: Oh, Yeongtak, et al.
Published: (2024)
Contextualized Visual Personalization in Vision-Language Models
by: Oh, Yeongtak, et al.
Published: (2026)
by: Oh, Yeongtak, et al.
Published: (2026)
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
by: Park, Sangha, et al.
Published: (2025)
by: Park, Sangha, et al.
Published: (2025)
ControlDreamer: Blending Geometry and Style in Text-to-3D
by: Oh, Yeongtak, et al.
Published: (2023)
by: Oh, Yeongtak, et al.
Published: (2023)
RePIC: Reinforced Post-Training for Personalizing Multi-Modal Language Models
by: Oh, Yeongtak, et al.
Published: (2025)
by: Oh, Yeongtak, et al.
Published: (2025)
Efficient Diffusion-Driven Corruption Editor for Test-Time Adaptation
by: Oh, Yeongtak, et al.
Published: (2024)
by: Oh, Yeongtak, et al.
Published: (2024)
Large-Scale Text-to-Image Model with Inpainting is a Zero-Shot Subject-Driven Image Generator
by: Shin, Chaehun, et al.
Published: (2024)
by: Shin, Chaehun, et al.
Published: (2024)
Improving Geometry in Sparse-View 3DGS via Reprojection-based DoF Separation
by: Kim, Yongsung, et al.
Published: (2024)
by: Kim, Yongsung, et al.
Published: (2024)
Omni-o3: Deep Nested Omnimodal Deduction for Deliberative Audio-Visual Reasoning
by: Zhang, Zhicheng, et al.
Published: (2026)
by: Zhang, Zhicheng, et al.
Published: (2026)
Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
by: Zhong, Hao, et al.
Published: (2025)
by: Zhong, Hao, et al.
Published: (2025)
OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
by: Tao, Keda, et al.
Published: (2025)
by: Tao, Keda, et al.
Published: (2025)
CKNN: Cleansed k-Nearest Neighbor for Unsupervised Video Anomaly Detection
by: Yi, Jihun, et al.
Published: (2024)
by: Yi, Jihun, et al.
Published: (2024)
Interactive Text-to-Image Retrieval with Large Language Models: A Plug-and-Play Approach
by: Lee, Saehyung, et al.
Published: (2024)
by: Lee, Saehyung, et al.
Published: (2024)
ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation
by: Liang, Yongyuan, et al.
Published: (2025)
by: Liang, Yongyuan, et al.
Published: (2025)
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models
by: Jung, Mingi, et al.
Published: (2025)
by: Jung, Mingi, et al.
Published: (2025)
LiViBench: An Omnimodal Benchmark for Interactive Livestream Video Understanding
by: Wang, Xiaodong, et al.
Published: (2026)
by: Wang, Xiaodong, et al.
Published: (2026)
Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage
by: Lee, Saehyung, et al.
Published: (2024)
by: Lee, Saehyung, et al.
Published: (2024)
Improving Diffusion-Based Generative Models via Approximated Optimal Transport
by: Kim, Daegyu, et al.
Published: (2024)
by: Kim, Daegyu, et al.
Published: (2024)
Normality Addition via Normality Detection in Industrial Image Anomaly Detection Models
by: Yi, Jihun, et al.
Published: (2024)
by: Yi, Jihun, et al.
Published: (2024)
DCText: Scheduled Attention Masking for Visual Text Generation via Divide-and-Conquer Strategy
by: Song, Jaewoo, et al.
Published: (2025)
by: Song, Jaewoo, et al.
Published: (2025)
Textual Training for the Hassle-Free Removal of Unwanted Visual Data: Case Studies on OOD and Hateful Image Detection
by: Lee, Saehyung, et al.
Published: (2024)
by: Lee, Saehyung, et al.
Published: (2024)
STAG: Structural Test-time Alignment of Gradients for Online Adaptation
by: Shin, Juhyeon, et al.
Published: (2024)
by: Shin, Juhyeon, et al.
Published: (2024)
HeSS: Head Sensitivity Score for Sparsity Redistribution in VGGT
by: Kim, Yongsung, et al.
Published: (2026)
by: Kim, Yongsung, et al.
Published: (2026)
SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
by: Park, Sangha, et al.
Published: (2025)
by: Park, Sangha, et al.
Published: (2025)
PersonaBooth: Personalized Text-to-Motion Generation
by: Kim, Boeun, et al.
Published: (2025)
by: Kim, Boeun, et al.
Published: (2025)
PhoStream: Benchmarking Real-World Streaming for Omnimodal Assistants in Mobile Scenarios
by: Lu, Xudong, et al.
Published: (2026)
by: Lu, Xudong, et al.
Published: (2026)
Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP
by: Park, Junsung, et al.
Published: (2025)
by: Park, Junsung, et al.
Published: (2025)
Diagnosing and Correcting Concept Omission in Multimodal Diffusion Transformers
by: Baek, Kanghyun, et al.
Published: (2026)
by: Baek, Kanghyun, et al.
Published: (2026)
CALM-Net: Curvature-Aware LiDAR Point Cloud-based Multi-Branch Neural Network for Vehicle Re-Identification
by: Lee, Dongwook, et al.
Published: (2025)
by: Lee, Dongwook, et al.
Published: (2025)
OmniLight: One Model to Rule All Lighting Conditions
by: Oh, Youngjin, et al.
Published: (2026)
by: Oh, Youngjin, et al.
Published: (2026)
Domain Generalization for Person Re-identification: A Survey Towards Domain-Agnostic Person Matching
by: Lee, Hyeonseo, et al.
Published: (2025)
by: Lee, Hyeonseo, et al.
Published: (2025)
EgoSocial: Benchmarking Proactive Intervention Ability of Omnimodal LLMs via Egocentric Social Interaction Perception
by: Wang, Xijun, et al.
Published: (2025)
by: Wang, Xijun, et al.
Published: (2025)
Active Perception Agent for Omnimodal Audio-Video Understanding
by: Tao, Keda, et al.
Published: (2025)
by: Tao, Keda, et al.
Published: (2025)
PersonaCraft: Personalized and Controllable Full-Body Multi-Human Scene Generation Using Occlusion-Aware 3D-Conditioned Diffusion
by: Kim, Gwanghyun, et al.
Published: (2024)
by: Kim, Gwanghyun, et al.
Published: (2024)
CMDA: Cross-Modal and Domain Adversarial Adaptation for LiDAR-Based 3D Object Detection
by: Chang, Gyusam, et al.
Published: (2024)
by: Chang, Gyusam, et al.
Published: (2024)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
by: Kim, Eunji, et al.
Published: (2024)
by: Kim, Eunji, et al.
Published: (2024)
Exploring the Potential of LLMs as Personalized Assistants: Dataset, Evaluation, and Analysis
by: Mok, Jisoo, et al.
Published: (2025)
by: Mok, Jisoo, et al.
Published: (2025)
OpenOmni: Advancing Open-Source Omnimodal Large Language Models with Progressive Multimodal Alignment and Real-Time Self-Aware Emotional Speech Synthesis
by: Luo, Run, et al.
Published: (2025)
by: Luo, Run, et al.
Published: (2025)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
by: Lew, Jaihyun, et al.
Published: (2024)
by: Lew, Jaihyun, et al.
Published: (2024)
Similar Items
-
Style-Friendly SNR Sampler for Style-Driven Generation
by: Choi, Jooyoung, et al.
Published: (2024) -
On mitigating stability-plasticity dilemma in CLIP-guided image morphing via geodesic distillation loss
by: Oh, Yeongtak, et al.
Published: (2024) -
Contextualized Visual Personalization in Vision-Language Models
by: Oh, Yeongtak, et al.
Published: (2026) -
Guiding What Not to Generate: Automated Negative Prompting for Text-Image Alignment
by: Park, Sangha, et al.
Published: (2025) -
ControlDreamer: Blending Geometry and Style in Text-to-3D
by: Oh, Yeongtak, et al.
Published: (2023)