DaMO: A Data-Efficient Multimodal Orchestrator for Temporal Reasoning with Video LLMs
Fuente:
arXiv
Guardado en:
| Autores principales: | Chiu, Bo-Cheng, Chen, Jen-Jee, Tseng, Yu-Chee, Chen, Feng-Chi, Yen, An-Zi |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Learning Displacement-Aware WiFi Representations for Weakly Supervised Relative Localization
por: Wei, Tzu-Ti, et al.
Publicado: (2026)
por: Wei, Tzu-Ti, et al.
Publicado: (2026)
Confidence-Driven Multi-Scale Model Selection for Cost-Efficient Inference
por: Chen, Bo-Wei, et al.
Publicado: (2026)
por: Chen, Bo-Wei, et al.
Publicado: (2026)
WiFi2Cap: Semantic Action Captioning from Wi-Fi CSI via Limb-Level Semantic Alignment
por: Wei, Tzu-Ti, et al.
Publicado: (2026)
por: Wei, Tzu-Ti, et al.
Publicado: (2026)
Refining Financial Consumer Complaints through Multi-Scale Model Interaction
por: Chen, Bo-Wei, et al.
Publicado: (2025)
por: Chen, Bo-Wei, et al.
Publicado: (2025)
ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos
por: Chen, Jr-Jen, et al.
Publicado: (2024)
por: Chen, Jr-Jen, et al.
Publicado: (2024)
Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video Understanding
por: Hu, Pengfei, et al.
Publicado: (2025)
por: Hu, Pengfei, et al.
Publicado: (2025)
TemporalVLM: Video LLMs for Temporal Reasoning in Long Videos
por: Fateh, Fawad Javed, et al.
Publicado: (2024)
por: Fateh, Fawad Javed, et al.
Publicado: (2024)
Text Slider: Efficient and Plug-and-Play Continuous Concept Control for Image/Video Synthesis via LoRA Adapters
por: Chiu, Pin-Yen, et al.
Publicado: (2025)
por: Chiu, Pin-Yen, et al.
Publicado: (2025)
Invisible Backdoor Triggers in Image Editing Model via Deep Watermarking
por: Chen, Yu-Feng, et al.
Publicado: (2025)
por: Chen, Yu-Feng, et al.
Publicado: (2025)
LingBench++: A Linguistically-Informed Benchmark and Reasoning Framework for Multi-Step and Cross-Cultural Inference with LLMs
por: Lian, Da-Chen, et al.
Publicado: (2025)
por: Lian, Da-Chen, et al.
Publicado: (2025)
V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning
por: Cheng, Zixu, et al.
Publicado: (2025)
por: Cheng, Zixu, et al.
Publicado: (2025)
Can Multimodal LLMs See Science Instruction? Benchmarking Pedagogical Reasoning in K-12 Classroom Videos
por: Shen, Yixuan, et al.
Publicado: (2026)
por: Shen, Yixuan, et al.
Publicado: (2026)
ISSR: Iterative Selection with Self-Review for Vocabulary Test Distractor Generation
por: Liu, Yu-Cheng, et al.
Publicado: (2025)
por: Liu, Yu-Cheng, et al.
Publicado: (2025)
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
por: Imam, Mohamed Fazli, et al.
Publicado: (2025)
por: Imam, Mohamed Fazli, et al.
Publicado: (2025)
Octopus: Agentic Multimodal Reasoning with Six-Capability Orchestration
por: Guo, Yifu, et al.
Publicado: (2025)
por: Guo, Yifu, et al.
Publicado: (2025)
Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
por: Shu, Yan, et al.
Publicado: (2025)
por: Shu, Yan, et al.
Publicado: (2025)
Learning-Based WiFi Fingerprint Inpainting via Generative Adversarial Networks
por: Chan, Yu, et al.
Publicado: (2024)
por: Chan, Yu, et al.
Publicado: (2024)
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models
por: Liang, Yiming, et al.
Publicado: (2026)
por: Liang, Yiming, et al.
Publicado: (2026)
TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
por: Zhang, Jun, et al.
Publicado: (2025)
por: Zhang, Jun, et al.
Publicado: (2025)
R-AVST: Empowering Video-LLMs with Fine-Grained Spatio-Temporal Reasoning in Complex Audio-Visual Scenarios
por: Zhu, Lu, et al.
Publicado: (2025)
por: Zhu, Lu, et al.
Publicado: (2025)
How Should Video LLMs Output Time? An Analysis of Efficient Temporal Grounding Paradigms
por: Jin, Shengji, et al.
Publicado: (2026)
por: Jin, Shengji, et al.
Publicado: (2026)
Paraphrase-Aligned Machine Translation
por: Chang, Ke-Ching, et al.
Publicado: (2024)
por: Chang, Ke-Ching, et al.
Publicado: (2024)
Language Matters: How Do Multilingual Input and Reasoning Paths Affect Large Reasoning Models?
por: Tam, Zhi Rui, et al.
Publicado: (2025)
por: Tam, Zhi Rui, et al.
Publicado: (2025)
Enhancing Temporal Modeling of Video LLMs via Time Gating
por: Hu, Zi-Yuan, et al.
Publicado: (2024)
por: Hu, Zi-Yuan, et al.
Publicado: (2024)
STORM: Token-Efficient Long Video Understanding for Multimodal LLMs
por: Jiang, Jindong, et al.
Publicado: (2025)
por: Jiang, Jindong, et al.
Publicado: (2025)
AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
por: Chen, Xinlong, et al.
Publicado: (2025)
por: Chen, Xinlong, et al.
Publicado: (2025)
VISTA: A Generative Egocentric Video Framework for Daily Assistance
por: Liu, Yu-Hsiang, et al.
Publicado: (2026)
por: Liu, Yu-Hsiang, et al.
Publicado: (2026)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
por: Pramanick, Shraman, et al.
Publicado: (2025)
por: Pramanick, Shraman, et al.
Publicado: (2025)
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
por: Song, Mingyang, et al.
Publicado: (2026)
por: Song, Mingyang, et al.
Publicado: (2026)
Quantum-Enhanced Temporal Embeddings via a Hybrid Seq2Seq Architecture
por: Hsieh, Tien-Ching, et al.
Publicado: (2026)
por: Hsieh, Tien-Ching, et al.
Publicado: (2026)
CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion
por: Yu, Shoubin, et al.
Publicado: (2024)
por: Yu, Shoubin, et al.
Publicado: (2024)
SafePLUG: Empowering Multimodal LLMs with Pixel-Level Insight and Temporal Grounding for Traffic Accident Understanding
por: Sheng, Zihao, et al.
Publicado: (2025)
por: Sheng, Zihao, et al.
Publicado: (2025)
GaMO: Geometry-aware Multi-view Diffusion Outpainting for Sparse-View 3D Reconstruction
por: Huang, Yi-Chuan, et al.
Publicado: (2025)
por: Huang, Yi-Chuan, et al.
Publicado: (2025)
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
por: Ma, David, et al.
Publicado: (2025)
por: Ma, David, et al.
Publicado: (2025)
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
por: Cheng, Zesen, et al.
Publicado: (2024)
por: Cheng, Zesen, et al.
Publicado: (2024)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
por: Jiang, Ruixiang, et al.
Publicado: (2025)
por: Jiang, Ruixiang, et al.
Publicado: (2025)
Stay Hungry, Stay Foolish: On the Extended Reading Articles Generation with LLMs
por: Liou, Yow-Fu, et al.
Publicado: (2025)
por: Liou, Yow-Fu, et al.
Publicado: (2025)
Thinking with Video: Video Generation as a Promising Multimodal Reasoning Paradigm
por: Tong, Jingqi, et al.
Publicado: (2025)
por: Tong, Jingqi, et al.
Publicado: (2025)
IntenBot: Flexible and Imprecise Multimodal Input for LLMs to Understand User Intentions for Casual and Human-Like HRI
por: Liu, Yen-Ting, et al.
Publicado: (2026)
por: Liu, Yen-Ting, et al.
Publicado: (2026)
Unlocking Exocentric Video-Language Data for Egocentric Video Representation Learning
por: Dou, Zi-Yi, et al.
Publicado: (2024)
por: Dou, Zi-Yi, et al.
Publicado: (2024)
Ejemplares similares
-
Learning Displacement-Aware WiFi Representations for Weakly Supervised Relative Localization
por: Wei, Tzu-Ti, et al.
Publicado: (2026) -
Confidence-Driven Multi-Scale Model Selection for Cost-Efficient Inference
por: Chen, Bo-Wei, et al.
Publicado: (2026) -
WiFi2Cap: Semantic Action Captioning from Wi-Fi CSI via Limb-Level Semantic Alignment
por: Wei, Tzu-Ti, et al.
Publicado: (2026) -
Refining Financial Consumer Complaints through Multi-Scale Model Interaction
por: Chen, Bo-Wei, et al.
Publicado: (2025) -
ReXTime: A Benchmark Suite for Reasoning-Across-Time in Videos
por: Chen, Jr-Jen, et al.
Publicado: (2024)