Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Ge, Yunhao, Zeng, Xiaohui, Huffman, Jacob Samuel, Lin, Tsung-Yi, Liu, Ming-Yu, Cui, Yin |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Describe Anything: Detailed Localized Image and Video Captioning
di: Lian, Long, et al.
Pubblicazione: (2025)
di: Lian, Long, et al.
Pubblicazione: (2025)
Meshtron: High-Fidelity, Artist-Like 3D Mesh Generation at Scale
di: Hao, Zekun, et al.
Pubblicazione: (2024)
di: Hao, Zekun, et al.
Pubblicazione: (2024)
Edify 3D: Scalable High-Quality 3D Asset Generation
di: NVIDIA, et al.
Pubblicazione: (2024)
di: NVIDIA, et al.
Pubblicazione: (2024)
Generating Accurate and Detailed Captions for High-Resolution Images
di: Lee, Hankyeol, et al.
Pubblicazione: (2025)
di: Lee, Hankyeol, et al.
Pubblicazione: (2025)
Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation
di: Ling, Lu, et al.
Pubblicazione: (2025)
di: Ling, Lu, et al.
Pubblicazione: (2025)
FlashVideo: Flowing Fidelity to Detail for Efficient High-Resolution Video Generation
di: Zhang, Shilong, et al.
Pubblicazione: (2025)
di: Zhang, Shilong, et al.
Pubblicazione: (2025)
Harnessing Caption Detailness for Data-Efficient Text-to-Image Generation
di: Wang, Xinran, et al.
Pubblicazione: (2025)
di: Wang, Xinran, et al.
Pubblicazione: (2025)
Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning
di: Ye, Qinghao, et al.
Pubblicazione: (2025)
di: Ye, Qinghao, et al.
Pubblicazione: (2025)
Benchmarking and Improving Detail Image Caption
di: Dong, Hongyuan, et al.
Pubblicazione: (2024)
di: Dong, Hongyuan, et al.
Pubblicazione: (2024)
EvoMakeup: High-Fidelity and Controllable Makeup Editing with MakeupQuad
di: Wu, Huadong, et al.
Pubblicazione: (2025)
di: Wu, Huadong, et al.
Pubblicazione: (2025)
Spatially Prompted Visual Trajectory Prediction for Egocentric Manipulation
di: Li, Yifan, et al.
Pubblicazione: (2026)
di: Li, Yifan, et al.
Pubblicazione: (2026)
CodecCap: High-Fidelity Codec-Inspired Residual Modeling for Dense Video Captioning
di: Lin, Zihan, et al.
Pubblicazione: (2026)
di: Lin, Zihan, et al.
Pubblicazione: (2026)
GCC: Generative Color Constancy via Diffusing a Color Checker
di: Chang, Chen-Wei, et al.
Pubblicazione: (2025)
di: Chang, Chen-Wei, et al.
Pubblicazione: (2025)
Point-It-Out: Benchmarking Embodied Reasoning for Vision Language Models in Multi-Stage Visual Grounding
di: Xue, Haotian, et al.
Pubblicazione: (2025)
di: Xue, Haotian, et al.
Pubblicazione: (2025)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
di: Wu, Peiran, et al.
Pubblicazione: (2025)
di: Wu, Peiran, et al.
Pubblicazione: (2025)
SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning
di: Zhang, Lin, et al.
Pubblicazione: (2025)
di: Zhang, Lin, et al.
Pubblicazione: (2025)
FontCrafter: High-Fidelity Element-Driven Artistic Font Creation with Visual In-Context Generation
di: Luo, Wuyang, et al.
Pubblicazione: (2026)
di: Luo, Wuyang, et al.
Pubblicazione: (2026)
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
di: Lei, Zhenxin, et al.
Pubblicazione: (2025)
di: Lei, Zhenxin, et al.
Pubblicazione: (2025)
Edify Image: High-Quality Image Generation with Pixel Space Laplacian Diffusion Models
di: NVIDIA, et al.
Pubblicazione: (2024)
di: NVIDIA, et al.
Pubblicazione: (2024)
DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
di: Li, Xiangtai, et al.
Pubblicazione: (2025)
di: Li, Xiangtai, et al.
Pubblicazione: (2025)
WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)
di: Zhuang, Shaobin, et al.
Pubblicazione: (2025)
DuoGen: Towards General Purpose Interleaved Multimodal Generation
di: Shi, Min, et al.
Pubblicazione: (2026)
di: Shi, Min, et al.
Pubblicazione: (2026)
SuperCarver: Texture-Consistent 3D Geometry Super-Resolution for High-Fidelity Surface Detail Generation
di: Zhang, Qijian, et al.
Pubblicazione: (2025)
di: Zhang, Qijian, et al.
Pubblicazione: (2025)
good4cir: Generating Detailed Synthetic Captions for Composed Image Retrieval
di: Kolouju, Pranavi, et al.
Pubblicazione: (2025)
di: Kolouju, Pranavi, et al.
Pubblicazione: (2025)
ReflectCAP: Detailed Image Captioning with Reflective Memory
di: Min, Kyungmin, et al.
Pubblicazione: (2026)
di: Min, Kyungmin, et al.
Pubblicazione: (2026)
The Devil is in the EOS: Sequence Training for Detailed Image Captioning
di: Mohamed, Abdelrahman, et al.
Pubblicazione: (2025)
di: Mohamed, Abdelrahman, et al.
Pubblicazione: (2025)
Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption
di: Qin, Luozheng, et al.
Pubblicazione: (2025)
di: Qin, Luozheng, et al.
Pubblicazione: (2025)
HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images
di: Liu, Yichen, et al.
Pubblicazione: (2026)
di: Liu, Yichen, et al.
Pubblicazione: (2026)
Hunyuan3D 2.5: Towards High-Fidelity 3D Assets Generation with Ultimate Details
di: Lai, Zeqiang, et al.
Pubblicazione: (2025)
di: Lai, Zeqiang, et al.
Pubblicazione: (2025)
AtomoVideo: High Fidelity Image-to-Video Generation
di: Gong, Litong, et al.
Pubblicazione: (2024)
di: Gong, Litong, et al.
Pubblicazione: (2024)
ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary
di: Gu, Zeqi, et al.
Pubblicazione: (2025)
di: Gu, Zeqi, et al.
Pubblicazione: (2025)
HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation
di: Shi, Wangzheng, et al.
Pubblicazione: (2025)
di: Shi, Wangzheng, et al.
Pubblicazione: (2025)
Omni-Captioner: Data Pipeline, Models, and Benchmark for Omni Detailed Perception
di: Ma, Ziyang, et al.
Pubblicazione: (2025)
di: Ma, Ziyang, et al.
Pubblicazione: (2025)
OwlCap: Harmonizing Motion-Detail for Video Captioning via HMD-270K and Caption Set Equivalence Reward
di: Zhong, Chunlin, et al.
Pubblicazione: (2025)
di: Zhong, Chunlin, et al.
Pubblicazione: (2025)
FiDeSR: High-Fidelity and Detail-Preserving One-Step Diffusion Super-Resolution
di: Kim, Aro, et al.
Pubblicazione: (2026)
di: Kim, Aro, et al.
Pubblicazione: (2026)
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
di: Gaur, Manu, et al.
Pubblicazione: (2024)
di: Gaur, Manu, et al.
Pubblicazione: (2024)
AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
di: Chai, Wenhao, et al.
Pubblicazione: (2024)
di: Chai, Wenhao, et al.
Pubblicazione: (2024)
IAR2: Improving Autoregressive Visual Generation with Semantic-Detail Associated Token Prediction
di: Yi, Ran, et al.
Pubblicazione: (2025)
di: Yi, Ran, et al.
Pubblicazione: (2025)
Towards High-Fidelity 3D Portrait Generation with Rich Details by Cross-View Prior-Aware Diffusion
di: Wei, Haoran, et al.
Pubblicazione: (2024)
di: Wei, Haoran, et al.
Pubblicazione: (2024)
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
di: Song, Jiahe, et al.
Pubblicazione: (2025)
di: Song, Jiahe, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Describe Anything: Detailed Localized Image and Video Captioning
di: Lian, Long, et al.
Pubblicazione: (2025) -
Meshtron: High-Fidelity, Artist-Like 3D Mesh Generation at Scale
di: Hao, Zekun, et al.
Pubblicazione: (2024) -
Edify 3D: Scalable High-Quality 3D Asset Generation
di: NVIDIA, et al.
Pubblicazione: (2024) -
Generating Accurate and Detailed Captions for High-Resolution Images
di: Lee, Hankyeol, et al.
Pubblicazione: (2025) -
Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation
di: Ling, Lu, et al.
Pubblicazione: (2025)