Aligning Vision Models with Human Aesthetics in Retrieval: Benchmarks and Algorithms
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Miaosen, Wei, Yixuan, Xing, Zhen, Ma, Yifei, Wu, Zuxuan, Li, Ji, Zhang, Zheng, Dai, Qi, Luo, Chong, Geng, Xin, Guo, Baining |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
MageBench: Bridging Large Multimodal Models to Agents
by: Zhang, Miaosen, et al.
Published: (2024)
by: Zhang, Miaosen, et al.
Published: (2024)
Covering Human Action Space for Computer Use: Data Synthesis and Benchmark
by: Zhang, Miaosen, et al.
Published: (2026)
by: Zhang, Miaosen, et al.
Published: (2026)
Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM Training
by: Zhang, Miaosen, et al.
Published: (2026)
by: Zhang, Miaosen, et al.
Published: (2026)
StableAnimator: High-Quality Identity-Preserving Human Image Animation
by: Tu, Shuyuan, et al.
Published: (2024)
by: Tu, Shuyuan, et al.
Published: (2024)
StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation
by: Tu, Shuyuan, et al.
Published: (2025)
by: Tu, Shuyuan, et al.
Published: (2025)
Phi-Ground Tech Report: Advancing Perception in GUI Grounding
by: Zhang, Miaosen, et al.
Published: (2025)
by: Zhang, Miaosen, et al.
Published: (2025)
Understanding Retrieval-Augmented Task Adaptation for Vision-Language Models
by: Ming, Yifei, et al.
Published: (2024)
by: Ming, Yifei, et al.
Published: (2024)
Beauty in the Eye of AI: Aligning LLMs and Vision Models with Human Aesthetics in Network Visualization
by: Zhang, Peng, et al.
Published: (2026)
by: Zhang, Peng, et al.
Published: (2026)
InfoAgent: Advancing Autonomous Information-Seeking Agents
by: Zhang, Gongrui, et al.
Published: (2025)
by: Zhang, Gongrui, et al.
Published: (2025)
MagicMotion: Controllable Video Generation with Dense-to-Sparse Trajectory Guidance
by: Li, Quanhao, et al.
Published: (2025)
by: Li, Quanhao, et al.
Published: (2025)
Improved Noise Schedule for Diffusion Training
by: Hang, Tiankai, et al.
Published: (2024)
by: Hang, Tiankai, et al.
Published: (2024)
RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents
by: Zhu, Jialiang, et al.
Published: (2026)
by: Zhu, Jialiang, et al.
Published: (2026)
AID: Adapting Image2Video Diffusion Models for Instruction-guided Video Prediction
by: Xing, Zhen, et al.
Published: (2024)
by: Xing, Zhen, et al.
Published: (2024)
FlashPortrait: 6x Faster Infinite Portrait Animation with Adaptive Latent Prediction
by: Tu, Shuyuan, et al.
Published: (2025)
by: Tu, Shuyuan, et al.
Published: (2025)
RelationVLM: Making Large Vision-Language Models Understand Visual Relations
by: Huang, Zhipeng, et al.
Published: (2024)
by: Huang, Zhipeng, et al.
Published: (2024)
StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation
by: Tu, Shuyuan, et al.
Published: (2025)
by: Tu, Shuyuan, et al.
Published: (2025)
CiQi-Agent: Aligning Vision, Tools and Aesthetics in Multimodal Agent for Cultural Reasoning on Chinese Porcelains
by: Wang, Wenhan, et al.
Published: (2026)
by: Wang, Wenhan, et al.
Published: (2026)
CCA: Collaborative Competitive Agents for Image Editing
by: Hang, Tiankai, et al.
Published: (2024)
by: Hang, Tiankai, et al.
Published: (2024)
A$^3$: Towards Advertising Aesthetic Assessment
by: Ji, Kaiyuan, et al.
Published: (2026)
by: Ji, Kaiyuan, et al.
Published: (2026)
Mitigating Hallucination in Large Vision-Language Models through Aligning Attention Distribution to Information Flow
by: Zhao, Jianfei, et al.
Published: (2025)
by: Zhao, Jianfei, et al.
Published: (2025)
Evaluating Menu OCR and Translation: A Benchmark for Aligning Human and Automated Evaluations in Large Vision-Language Models
by: Wu, Zhanglin, et al.
Published: (2025)
by: Wu, Zhanglin, et al.
Published: (2025)
IRGen: Generative Modeling for Image Retrieval
by: Zhang, Yidan, et al.
Published: (2023)
by: Zhang, Yidan, et al.
Published: (2023)
VisualCritic: Making LMMs Perceive Visual Quality Like Humans
by: Huang, Zhipeng, et al.
Published: (2024)
by: Huang, Zhipeng, et al.
Published: (2024)
How Does Fine-Tuning Impact Out-of-Distribution Detection for Vision-Language Models?
by: Ming, Yifei, et al.
Published: (2023)
by: Ming, Yifei, et al.
Published: (2023)
DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation
by: Peng, Yuang, et al.
Published: (2024)
by: Peng, Yuang, et al.
Published: (2024)
When Text Hijacks Vision: Benchmarking and Mitigating Text Overlay-Induced Hallucination in Vision Language Models
by: Yakun, Cui, et al.
Published: (2026)
by: Yakun, Cui, et al.
Published: (2026)
FlashMotion: Few-Step Controllable Video Generation with Trajectory Guidance
by: Li, Quanhao, et al.
Published: (2026)
by: Li, Quanhao, et al.
Published: (2026)
Aligning Anime Video Generation with Human Feedback
by: Zhu, Bingwen, et al.
Published: (2025)
by: Zhu, Bingwen, et al.
Published: (2025)
Language-Guided Face Animation by Recurrent StyleGAN-based Generator
by: Hang, Tiankai, et al.
Published: (2022)
by: Hang, Tiankai, et al.
Published: (2022)
Towards Explainable Fusion and Balanced Learning in Multimodal Sentiment Analysis
by: Luo, Miaosen, et al.
Published: (2025)
by: Luo, Miaosen, et al.
Published: (2025)
ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws
by: Li, Ruihang, et al.
Published: (2024)
by: Li, Ruihang, et al.
Published: (2024)
AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
by: Zhou, Ziwei, et al.
Published: (2026)
by: Zhou, Ziwei, et al.
Published: (2026)
RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models
by: Yan, Qihang, et al.
Published: (2025)
by: Yan, Qihang, et al.
Published: (2025)
RoboReflect: A Robotic Reflective Reasoning Framework for Grasping Ambiguous-Condition Objects
by: Luo, Zhen, et al.
Published: (2025)
by: Luo, Zhen, et al.
Published: (2025)
FIOVA: A Multi-Annotator Benchmark for Human-Aligned Video Captioning
by: Hu, Shiyu, et al.
Published: (2024)
by: Hu, Shiyu, et al.
Published: (2024)
@Bench: Benchmarking Vision-Language Models for Human-centered Assistive Technology
by: Jiang, Xin, et al.
Published: (2024)
by: Jiang, Xin, et al.
Published: (2024)
MotionFollower: Editing Video Motion via Lightweight Score-Guided Diffusion
by: Tu, Shuyuan, et al.
Published: (2024)
by: Tu, Shuyuan, et al.
Published: (2024)
Re-Align: Aligning Vision Language Models via Retrieval-Augmented Direct Preference Optimization
by: Xing, Shuo, et al.
Published: (2025)
by: Xing, Shuo, et al.
Published: (2025)
AlignCoder: Aligning Retrieval with Target Intent for Repository-Level Code Completion
by: Jiang, Tianyue, et al.
Published: (2026)
by: Jiang, Tianyue, et al.
Published: (2026)
Pareto-Enhanced Portrait Generation: Vision-Aligned Text Supervision for Alignment, Realism, and Aesthetics
by: Wang, Yunlong, et al.
Published: (2026)
by: Wang, Yunlong, et al.
Published: (2026)
Similar Items
-
MageBench: Bridging Large Multimodal Models to Agents
by: Zhang, Miaosen, et al.
Published: (2024) -
Covering Human Action Space for Computer Use: Data Synthesis and Benchmark
by: Zhang, Miaosen, et al.
Published: (2026) -
Towards On-Policy SFT: Distribution Discriminant Theory and its Applications in LLM Training
by: Zhang, Miaosen, et al.
Published: (2026) -
StableAnimator: High-Quality Identity-Preserving Human Image Animation
by: Tu, Shuyuan, et al.
Published: (2024) -
StableAnimator++: Overcoming Pose Misalignment and Face Distortion for Human Image Animation
by: Tu, Shuyuan, et al.
Published: (2025)