Describe Anything: Detailed Localized Image and Video Captioning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lian, Long, Ding, Yifan, Ge, Yunhao, Liu, Sifei, Mao, Hanzi, Li, Boyi, Pavone, Marco, Liu, Ming-Yu, Darrell, Trevor, Yala, Adam, Cui, Yin |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
von: Lian, Long, et al.
Veröffentlicht: (2023)
von: Lian, Long, et al.
Veröffentlicht: (2023)
LLM-grounded Video Diffusion Models
von: Lian, Long, et al.
Veröffentlicht: (2023)
von: Lian, Long, et al.
Veröffentlicht: (2023)
Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation
von: Ge, Yunhao, et al.
Veröffentlicht: (2024)
von: Ge, Yunhao, et al.
Veröffentlicht: (2024)
Atlas: Multi-Scale Attention Improves Long Context Image Modeling
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2025)
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2025)
Scaling Vision Pre-Training to 4K Resolution
von: Shi, Baifeng, et al.
Veröffentlicht: (2025)
von: Shi, Baifeng, et al.
Veröffentlicht: (2025)
FlexCap: Describe Anything in Images in Controllable Detail
von: Dwibedi, Debidatta, et al.
Veröffentlicht: (2024)
von: Dwibedi, Debidatta, et al.
Veröffentlicht: (2024)
Wolf: Dense Video Captioning with a World Summarization Framework
von: Li, Boyi, et al.
Veröffentlicht: (2024)
von: Li, Boyi, et al.
Veröffentlicht: (2024)
Learning Adaptive Parallel Reasoning with Language Models
von: Pan, Jiayi, et al.
Veröffentlicht: (2025)
von: Pan, Jiayi, et al.
Veröffentlicht: (2025)
Segment Anything without Supervision
von: Wang, XuDong, et al.
Veröffentlicht: (2024)
von: Wang, XuDong, et al.
Veröffentlicht: (2024)
Separate Anything You Describe
von: Liu, Xubo, et al.
Veröffentlicht: (2023)
von: Liu, Xubo, et al.
Veröffentlicht: (2023)
Rethinking Patch Dependence for Masked Autoencoders
von: Fu, Letian, et al.
Veröffentlicht: (2024)
von: Fu, Letian, et al.
Veröffentlicht: (2024)
UGC-VideoCaptioner: An Omni UGC Video Detail Caption Model and New Benchmarks
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
von: Wu, Peiran, et al.
Veröffentlicht: (2025)
TULIP: Towards Unified Language-Image Pretraining
von: Tang, Zineng, et al.
Veröffentlicht: (2025)
von: Tang, Zineng, et al.
Veröffentlicht: (2025)
FoundationMotion: Auto-Labeling and Reasoning about Spatial Movement in Videos
von: Gan, Yulu, et al.
Veröffentlicht: (2025)
von: Gan, Yulu, et al.
Veröffentlicht: (2025)
Describe Anything in Medical Images
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
von: Xiao, Xi, et al.
Veröffentlicht: (2025)
LOTUS: A Leaderboard for Detailed Image Captioning from Quality to Societal Bias and User Preferences
von: Hirota, Yusuke, et al.
Veröffentlicht: (2025)
von: Hirota, Yusuke, et al.
Veröffentlicht: (2025)
Segment and Caption Anything
von: Huang, Xiaoke, et al.
Veröffentlicht: (2023)
von: Huang, Xiaoke, et al.
Veröffentlicht: (2023)
Caption Anything in Video: Fine-grained Object-centric Captioning via Spatiotemporal Multimodal Prompting
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)
von: Tang, Yunlong, et al.
Veröffentlicht: (2025)
UnSAMv2: Self-Supervised Learning Enables Segment Anything at Any Granularity
von: Yu, Junwei, et al.
Veröffentlicht: (2025)
von: Yu, Junwei, et al.
Veröffentlicht: (2025)
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos
von: Lin, Weifeng, et al.
Veröffentlicht: (2025)
von: Lin, Weifeng, et al.
Veröffentlicht: (2025)
DescribeEarth: Describe Anything for Remote Sensing Images
von: Li, Kaiyu, et al.
Veröffentlicht: (2025)
von: Li, Kaiyu, et al.
Veröffentlicht: (2025)
ThreadWeaver: Adaptive Threading for Efficient Parallel Reasoning in Language Models
von: Lian, Long, et al.
Veröffentlicht: (2025)
von: Lian, Long, et al.
Veröffentlicht: (2025)
Describe Anything Anywhere At Any Moment
von: Gorlo, Nicolas, et al.
Veröffentlicht: (2025)
von: Gorlo, Nicolas, et al.
Veröffentlicht: (2025)
Attend Before Attention: Efficient and Scalable Video Understanding via Autoregressive Gazing
von: Shi, Baifeng, et al.
Veröffentlicht: (2026)
von: Shi, Baifeng, et al.
Veröffentlicht: (2026)
Implicit Location-Caption Alignment via Complementary Masking for Weakly-Supervised Dense Video Captioning
von: Ge, Shiping, et al.
Veröffentlicht: (2024)
von: Ge, Shiping, et al.
Veröffentlicht: (2024)
Mind the Generative Details: Direct Localized Detail Preference Optimization for Video Diffusion Models
von: Huang, Zitong, et al.
Veröffentlicht: (2026)
von: Huang, Zitong, et al.
Veröffentlicht: (2026)
On a Mathematical Model Describing Chemotherapeutic Drug Treatment for Tumor Cells
von: Liu, Xiaoqin, et al.
Veröffentlicht: (2026)
von: Liu, Xiaoqin, et al.
Veröffentlicht: (2026)
CLAIR-A: Leveraging Large Language Models to Judge Audio Captions
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
von: Wu, Tsung-Han, et al.
Veröffentlicht: (2024)
Describing Differences in Image Sets with Natural Language
von: Dunlap, Lisa, et al.
Veröffentlicht: (2023)
von: Dunlap, Lisa, et al.
Veröffentlicht: (2023)
ArtiScene: Language-Driven Artistic 3D Scene Generation Through Image Intermediary
von: Gu, Zeqi, et al.
Veröffentlicht: (2025)
von: Gu, Zeqi, et al.
Veröffentlicht: (2025)
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details
von: Zhou, Dewei, et al.
Veröffentlicht: (2026)
von: Zhou, Dewei, et al.
Veröffentlicht: (2026)
Co-persona: Leveraging LLMs and Expert Collaboration to Understand User Personas through Social Media Data Analysis
von: Yin, Min, et al.
Veröffentlicht: (2025)
von: Yin, Min, et al.
Veröffentlicht: (2025)
Pillar-0: A New Frontier for Radiology Foundation Models
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2025)
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2025)
Cockatiel: Ensembling Synthetic and Human Preferenced Training for Detailed Video Caption
von: Qin, Luozheng, et al.
Veröffentlicht: (2025)
von: Qin, Luozheng, et al.
Veröffentlicht: (2025)
SynPO: Synergizing Descriptiveness and Preference Optimization for Video Detailed Captioning
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
von: Dang, Jisheng, et al.
Veröffentlicht: (2025)
URECA: Unique Region Caption Anything
von: Lim, Sangbeom, et al.
Veröffentlicht: (2025)
von: Lim, Sangbeom, et al.
Veröffentlicht: (2025)
Scalable Policy Evaluation with Video World Models
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
von: Tseng, Wei-Cheng, et al.
Veröffentlicht: (2025)
Hallucination Localization in Video Captioning
von: Nakada, Shota, et al.
Veröffentlicht: (2025)
von: Nakada, Shota, et al.
Veröffentlicht: (2025)
LocalStyleFool: Regional Video Style Transfer Attack Using Segment Anything Model
von: Cao, Yuxin, et al.
Veröffentlicht: (2024)
von: Cao, Yuxin, et al.
Veröffentlicht: (2024)
GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration
von: Xu, Wan, et al.
Veröffentlicht: (2025)
von: Xu, Wan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
von: Lian, Long, et al.
Veröffentlicht: (2023) -
LLM-grounded Video Diffusion Models
von: Lian, Long, et al.
Veröffentlicht: (2023) -
Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation
von: Ge, Yunhao, et al.
Veröffentlicht: (2024) -
Atlas: Multi-Scale Attention Improves Long Context Image Modeling
von: Agrawal, Kumar Krishna, et al.
Veröffentlicht: (2025) -
Scaling Vision Pre-Training to 4K Resolution
von: Shi, Baifeng, et al.
Veröffentlicht: (2025)