Chrono: A Simple Blueprint for Representing Time in MLLMs
Fuente:
arXiv
Saved in:
| Main Authors: | Rodriguez, Hector, Meinardus, Boris, Batra, Anil, Rohrbach, Anna, Rohrbach, Marcus |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
by: Abdessaied, Adnen, et al.
Published: (2025)
by: Abdessaied, Adnen, et al.
Published: (2025)
Predicting Implicit Arguments in Procedural Video Instructions
by: Batra, Anil, et al.
Published: (2025)
by: Batra, Anil, et al.
Published: (2025)
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
by: Braun, Tobias, et al.
Published: (2024)
by: Braun, Tobias, et al.
Published: (2024)
SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring
by: Rodriguez, Hector G., et al.
Published: (2026)
by: Rodriguez, Hector G., et al.
Published: (2026)
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
by: Rothermel, Mark, et al.
Published: (2026)
by: Rothermel, Mark, et al.
Published: (2026)
Efficient Pre-training for Localized Instruction Generation of Videos
by: Batra, Anil, et al.
Published: (2023)
by: Batra, Anil, et al.
Published: (2023)
HaloProbe: Bayesian Detection and Mitigation of Object Hallucinations in Vision-Language Models
by: Zohrabi, Reihaneh, et al.
Published: (2026)
by: Zohrabi, Reihaneh, et al.
Published: (2026)
Spurious-Aware Prototype Refinement for Reliable Out-of-Distribution Detection
by: Zohrabi, Reihaneh, et al.
Published: (2025)
by: Zohrabi, Reihaneh, et al.
Published: (2025)
ReCap: Lightweight Referential Grounding for Coherent Story Visualization
by: Arora, Aditya, et al.
Published: (2026)
by: Arora, Aditya, et al.
Published: (2026)
Diffusion Classifiers Understand Compositionality, but Conditions Apply
by: Jeong, Yujin, et al.
Published: (2025)
by: Jeong, Yujin, et al.
Published: (2025)
Tuning Just Enough: Lightweight Backdoor Attacks on Multi-Encoder Diffusion Models
by: Chen, Ziyuan, et al.
Published: (2026)
by: Chen, Ziyuan, et al.
Published: (2026)
Variational Visual Question Answering for Uncertainty-Aware Selective Prediction
by: Wieczorek, Tobias Jan, et al.
Published: (2025)
by: Wieczorek, Tobias Jan, et al.
Published: (2025)
When Do Diffusion Models learn to Generate Multiple Objects?
by: Jeong, Yujin, et al.
Published: (2026)
by: Jeong, Yujin, et al.
Published: (2026)
Evaluating the Impact of Post-Training Quantization on Reliable VQA with Multimodal LLMs
by: Kurz, Paul Jonas, et al.
Published: (2026)
by: Kurz, Paul Jonas, et al.
Published: (2026)
Multi-axis Analysis of Image Manipulation Localization
by: Nichols, Keanu, et al.
Published: (2026)
by: Nichols, Keanu, et al.
Published: (2026)
CAST: Cross-modal Alignment Similarity Test for Vision Language Models
by: Dagan, Gautier, et al.
Published: (2024)
by: Dagan, Gautier, et al.
Published: (2024)
Shape-Guided Diffusion with Inside-Outside Attention
by: Park, Dong Huk, et al.
Published: (2022)
by: Park, Dong Huk, et al.
Published: (2022)
ChronoMagic-Bench: A Benchmark for Metamorphic Evaluation of Text-to-Time-lapse Video Generation
by: Yuan, Shenghai, et al.
Published: (2024)
by: Yuan, Shenghai, et al.
Published: (2024)
HydroChronos: Forecasting Decades of Surface Water Change
by: Cambrin, Daniele Rege, et al.
Published: (2025)
by: Cambrin, Daniele Rege, et al.
Published: (2025)
SimpleOCR: Rendering Visualized Questions to Teach MLLMs to Read
by: Peng, Yibo, et al.
Published: (2026)
by: Peng, Yibo, et al.
Published: (2026)
When MLLMs Meet Compression Distortion: A Coding Paradigm Tailored to MLLMs
by: Liu, Jinming, et al.
Published: (2025)
by: Liu, Jinming, et al.
Published: (2025)
ChronosObserver: Taming 4D World with Hyperspace Diffusion Sampling
by: Wang, Qisen, et al.
Published: (2025)
by: Wang, Qisen, et al.
Published: (2025)
ChronoEdit: Towards Temporal Reasoning for Image Editing and World Simulation
by: Wu, Jay Zhangjie, et al.
Published: (2025)
by: Wu, Jay Zhangjie, et al.
Published: (2025)
GTPred: Benchmarking MLLMs for Interpretable Geo-localization and Time-of-capture Prediction
by: Li, Jinnao, et al.
Published: (2026)
by: Li, Jinnao, et al.
Published: (2026)
ChronoSC: Task-Oriented Semantic Communication via Temporal-to-Color Encoding
by: Nguyen, Phuc H., et al.
Published: (2026)
by: Nguyen, Phuc H., et al.
Published: (2026)
ChronoTailor: Harnessing Attention Guidance for Fine-Grained Video Virtual Try-On
by: Wang, Jinjuan, et al.
Published: (2025)
by: Wang, Jinjuan, et al.
Published: (2025)
HABIT: Chrono-Synergia Robust Progressive Learning Framework for Composed Image Retrieval
by: Li, Zixu, et al.
Published: (2026)
by: Li, Zixu, et al.
Published: (2026)
Law of Vision Representation in MLLMs
by: Yang, Shijia, et al.
Published: (2024)
by: Yang, Shijia, et al.
Published: (2024)
Benchmarking Large and Small MLLMs
by: Feng, Xuelu, et al.
Published: (2025)
by: Feng, Xuelu, et al.
Published: (2025)
Baton: Explicit Semantic Blueprints for Joint Video-Audio Generation
by: Tu, Shuyuan, et al.
Published: (2026)
by: Tu, Shuyuan, et al.
Published: (2026)
LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts
by: Gani, Hanan, et al.
Published: (2023)
by: Gani, Hanan, et al.
Published: (2023)
ChronoEarth-492K: A Large Scale and Long Horizon Spatiotemporal Hyperspectral Earth Observation Dataset and Benchmark
by: Si, Haozhe, et al.
Published: (2026)
by: Si, Haozhe, et al.
Published: (2026)
BioBench: A Blueprint to Move Beyond ImageNet for Scientific ML Benchmarks
by: Stevens, Samuel
Published: (2025)
by: Stevens, Samuel
Published: (2025)
Automated Multi-level Preference for MLLMs
by: Zhang, Mengxi, et al.
Published: (2024)
by: Zhang, Mengxi, et al.
Published: (2024)
Training-Free Reasoning and Reflection in MLLMs
by: Wei, Hongchen, et al.
Published: (2025)
by: Wei, Hongchen, et al.
Published: (2025)
MokA: Multimodal Low-Rank Adaptation for MLLMs
by: Wei, Yake, et al.
Published: (2025)
by: Wei, Yake, et al.
Published: (2025)
RealCustom++: Representing Images as Real Textual Word for Real-Time Customization
by: Mao, Zhendong, et al.
Published: (2024)
by: Mao, Zhendong, et al.
Published: (2024)
Visual Jigsaw Post-Training Improves MLLMs
by: Wu, Penghao, et al.
Published: (2025)
by: Wu, Penghao, et al.
Published: (2025)
Sketch-in-Latents: Eliciting Unified Reasoning in MLLMs
by: Tong, Jintao, et al.
Published: (2025)
by: Tong, Jintao, et al.
Published: (2025)
Why MLLMs Struggle to Determine Object Orientations
by: Gopinath, Anju, et al.
Published: (2026)
by: Gopinath, Anju, et al.
Published: (2026)
Similar Items
-
V$^2$Dial: Unification of Video and Visual Dialog via Multimodal Experts
by: Abdessaied, Adnen, et al.
Published: (2025) -
Predicting Implicit Arguments in Procedural Video Instructions
by: Batra, Anil, et al.
Published: (2025) -
DEFAME: Dynamic Evidence-based FAct-checking with Multimodal Experts
by: Braun, Tobias, et al.
Published: (2024) -
SIEVES: Selective Prediction Generalizes through Visual Evidence Scoring
by: Rodriguez, Hector G., et al.
Published: (2026) -
VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
by: Rothermel, Mark, et al.
Published: (2026)