Detailed Object Description with Controllable Dimensions
Fuente:
arXiv
Saved in:
| Main Authors: | Wang, Xinran, Zhang, Haiwen, Li, Baoteng, Liang, Kongming, Sun, Hao, He, Zhongjiang, Ma, Zhanyu, Guo, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
by: Wang, Xinran, et al.
Published: (2024)
by: Wang, Xinran, et al.
Published: (2024)
Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation
by: Li, Baoteng, et al.
Published: (2026)
by: Li, Baoteng, et al.
Published: (2026)
Benchmarking Semantic Segmentation Models via Appearance and Geometry Attribute Editing
by: Yin, Zijin, et al.
Published: (2026)
by: Yin, Zijin, et al.
Published: (2026)
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
by: Wang, Xinran, et al.
Published: (2025)
by: Wang, Xinran, et al.
Published: (2025)
Evaluating Attribute Comprehension in Large Vision-Language Models
by: Zhang, Haiwen, et al.
Published: (2024)
by: Zhang, Haiwen, et al.
Published: (2024)
FairHuman: Boosting Hand and Face Quality in Human Image Generation with Minimum Potential Delay Fairness in Diffusion Models
by: Wang, Yuxuan, et al.
Published: (2025)
by: Wang, Yuxuan, et al.
Published: (2025)
OmniEraser: Remove Objects and Their Effects in Images with Paired Video-Frame Data
by: Wei, Runpu, et al.
Published: (2025)
by: Wei, Runpu, et al.
Published: (2025)
Benchmarking Segmentation Models with Mask-Preserved Attribute Editing
by: Yin, Zijin, et al.
Published: (2024)
by: Yin, Zijin, et al.
Published: (2024)
DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions
by: Wang, Xinran, et al.
Published: (2026)
by: Wang, Xinran, et al.
Published: (2026)
Geometric Image Editing via Effects-Sensitive In-Context Inpainting with Diffusion Transformers
by: Zhang, Shuo, et al.
Published: (2026)
by: Zhang, Shuo, et al.
Published: (2026)
Efficient Face Super-Resolution via Wavelet-based Feature Enhancement Network
by: Li, Wenjie, et al.
Published: (2024)
by: Li, Wenjie, et al.
Published: (2024)
ConMo: Controllable Motion Disentanglement and Recomposition for Zero-Shot Motion Transfer
by: Gao, Jiayi, et al.
Published: (2025)
by: Gao, Jiayi, et al.
Published: (2025)
CineTechBench: A Benchmark for Cinematographic Technique Understanding and Generation
by: Wang, Xinran, et al.
Published: (2025)
by: Wang, Xinran, et al.
Published: (2025)
DriveRX: A Vision-Language Reasoning Model for Cross-Task Autonomous Driving
by: Diao, Muxi, et al.
Published: (2025)
by: Diao, Muxi, et al.
Published: (2025)
PGP-SAM: Prototype-Guided Prompt Learning for Efficient Few-Shot Medical Image Segmentation
by: Yan, Zhonghao, et al.
Published: (2025)
by: Yan, Zhonghao, et al.
Published: (2025)
Hepato-LLaVA: An Expert MLLM with Sparse Topo-Pack Attention for Hepatocellular Pathology Analysis on Whole Slide Images
by: Yang, Yuxuan, et al.
Published: (2026)
by: Yang, Yuxuan, et al.
Published: (2026)
Toward Generalizable Forgery Detection and Reasoning
by: Gao, Yueying, et al.
Published: (2025)
by: Gao, Yueying, et al.
Published: (2025)
Generative Visual Chain-of-Thought for Image Editing
by: Yin, Zijin, et al.
Published: (2026)
by: Yin, Zijin, et al.
Published: (2026)
Polyp-E: Benchmarking the Robustness of Deep Segmentation Models via Polyp Editing
by: Wei, Runpu, et al.
Published: (2024)
by: Wei, Runpu, et al.
Published: (2024)
Boosting Robust AIGI Detection with LoRA-based Pairwise Training
by: Xia, Ruiyang, et al.
Published: (2026)
by: Xia, Ruiyang, et al.
Published: (2026)
Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement
by: Liang, Xinyue, et al.
Published: (2025)
by: Liang, Xinyue, et al.
Published: (2025)
Exploring Dynamic Transformer for Efficient Object Tracking
by: Zhu, Jiawen, et al.
Published: (2024)
by: Zhu, Jiawen, et al.
Published: (2024)
Controllable-Continuous Color Editing in Diffusion Model via Color Mapping
by: Yang, Yuqi, et al.
Published: (2025)
by: Yang, Yuqi, et al.
Published: (2025)
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
by: Dang, Jisheng, et al.
Published: (2025)
by: Dang, Jisheng, et al.
Published: (2025)
FourierSR: A Fourier Token-based Plugin for Efficient Image Super-Resolution
by: Li, Wenjie, et al.
Published: (2025)
by: Li, Wenjie, et al.
Published: (2025)
SpecGen: Neural Spectral BRDF Generation via Spectral-Spatial Tri-plane Aggregation
by: Jin, Zhenyu, et al.
Published: (2025)
by: Jin, Zhenyu, et al.
Published: (2025)
Measurement-Constrained Sampling for Text-Prompted Blind Face Restoration
by: Li, Wenjie, et al.
Published: (2025)
by: Li, Wenjie, et al.
Published: (2025)
Self-Supervised Selective-Guided Diffusion Model for Old-Photo Face Restoration
by: Li, Wenjie, et al.
Published: (2025)
by: Li, Wenjie, et al.
Published: (2025)
MedReasoner: Reinforcement Learning Drives Reasoning Grounding from Clinical Thought to Pixel-Level Precision
by: Yan, Zhonghao, et al.
Published: (2025)
by: Yan, Zhonghao, et al.
Published: (2025)
WAT: Online Video Understanding Needs Watching Before Thinking
by: Han, Zifan, et al.
Published: (2026)
by: Han, Zifan, et al.
Published: (2026)
ProTA: Probabilistic Token Aggregation for Text-Video Retrieval
by: Fang, Han, et al.
Published: (2024)
by: Fang, Han, et al.
Published: (2024)
Harnessing Joint Rain-/Detail-aware Representations to Eliminate Intricate Rains
by: Ran, Wu, et al.
Published: (2024)
by: Ran, Wu, et al.
Published: (2024)
NeRSP: Neural 3D Reconstruction for Reflective Objects with Sparse Polarized Images
by: Han, Yufei, et al.
Published: (2024)
by: Han, Yufei, et al.
Published: (2024)
Seeing Through the Rain: Resolving High-Frequency Conflicts in Deraining and Super-Resolution via Diffusion Guidance
by: Li, Wenjie, et al.
Published: (2025)
by: Li, Wenjie, et al.
Published: (2025)
PMNI: Pose-free Multi-view Normal Integration for Reflective and Textureless Surface Reconstruction
by: Pei, Mingzhi, et al.
Published: (2025)
by: Pei, Mingzhi, et al.
Published: (2025)
MGIMM: Multi-Granularity Instruction Multimodal Model for Attribute-Guided Remote Sensing Image Detailed Description
by: Yang, Cong, et al.
Published: (2024)
by: Yang, Cong, et al.
Published: (2024)
Benchmarking and Improving Detail Image Caption
by: Dong, Hongyuan, et al.
Published: (2024)
by: Dong, Hongyuan, et al.
Published: (2024)
Disentangle and denoise: Tackling context misalignment for video moment retrieval
by: Ma, Kaijing, et al.
Published: (2024)
by: Ma, Kaijing, et al.
Published: (2024)
RotatedMVPS: Multi-view Photometric Stereo with Rotated Natural Light
by: Yang, Songyun, et al.
Published: (2025)
by: Yang, Songyun, et al.
Published: (2025)
Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions
by: Pi, Renjie, et al.
Published: (2024)
by: Pi, Renjie, et al.
Published: (2024)
Similar Items
-
From Simple to Professional: A Combinatorial Controllable Image Captioning Agent
by: Wang, Xinran, et al.
Published: (2024) -
Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation
by: Li, Baoteng, et al.
Published: (2026) -
Benchmarking Semantic Segmentation Models via Appearance and Geometry Attribute Editing
by: Yin, Zijin, et al.
Published: (2026) -
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
by: Wang, Xinran, et al.
Published: (2025) -
Evaluating Attribute Comprehension in Large Vision-Language Models
by: Zhang, Haiwen, et al.
Published: (2024)