NeMo: Needle in a Montage for Video-Language Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Hu, Zi-Yuan, Liang, Shuo, Zheng, Duo, Li, Yanyang, Tao, Yeyao, Huang, Shijia, Feng, Wei, Qin, Jia, Yu, Jianguang, Huang, Jing, Fang, Meng, Li, Yin, Wang, Liwei |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
by: Zheng, Duo, et al.
Published: (2025)
by: Zheng, Duo, et al.
Published: (2025)
Efficient-VLN: A Training-Efficient Vision-Language Navigation Model
by: Zheng, Duo, et al.
Published: (2025)
by: Zheng, Duo, et al.
Published: (2025)
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
by: Zheng, Duo, et al.
Published: (2024)
by: Zheng, Duo, et al.
Published: (2024)
Training Video Foundation Models with NVIDIA NeMo
by: Patel, Zeeshan, et al.
Published: (2025)
by: Patel, Zeeshan, et al.
Published: (2025)
Fine-grained Spatiotemporal Grounding on Egocentric Videos
by: Liang, Shuo, et al.
Published: (2025)
by: Liang, Shuo, et al.
Published: (2025)
NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment
by: Shen, Gerald, et al.
Published: (2024)
by: Shen, Gerald, et al.
Published: (2024)
NeMo-Inspector: A Visualization Tool for LLM Generation Analysis
by: Gitman, Daria, et al.
Published: (2025)
by: Gitman, Daria, et al.
Published: (2025)
X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention
by: Zhao, Xiaochen, et al.
Published: (2025)
by: Zhao, Xiaochen, et al.
Published: (2025)
Finding NeMo: Negative-mined Mosaic Augmentation for Referring Image Segmentation
by: Ha, Seongsu, et al.
Published: (2024)
by: Ha, Seongsu, et al.
Published: (2024)
Enhancing Temporal Modeling of Video LLMs via Time Gating
by: Hu, Zi-Yuan, et al.
Published: (2024)
by: Hu, Zi-Yuan, et al.
Published: (2024)
Making Long-Context Language Models Better Multi-Hop Reasoners
by: Li, Yanyang, et al.
Published: (2024)
by: Li, Yanyang, et al.
Published: (2024)
Towards Learning a Generalist Model for Embodied Navigation
by: Zheng, Duo, et al.
Published: (2023)
by: Zheng, Duo, et al.
Published: (2023)
Text-Video Multi-Grained Integration for Video Moment Montage
by: Yin, Zhihui, et al.
Published: (2024)
by: Yin, Zhihui, et al.
Published: (2024)
DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation
by: Liu, Jiawei, et al.
Published: (2025)
by: Liu, Jiawei, et al.
Published: (2025)
Learning to Reason from Feedback at Test-Time
by: Li, Yanyang, et al.
Published: (2025)
by: Li, Yanyang, et al.
Published: (2025)
Rethinking Chain-of-Thought Reasoning for Videos
by: Zhong, Yiwu, et al.
Published: (2025)
by: Zhong, Yiwu, et al.
Published: (2025)
Logic of Montage
by: Takahashi, Hayami, et al.
Published: (2025)
by: Takahashi, Hayami, et al.
Published: (2025)
Two Causally Related Needles in a Video Haystack
by: Li, Miaoyu, et al.
Published: (2025)
by: Li, Miaoyu, et al.
Published: (2025)
Enforced Interface Constraints for Domain Decomposition Method of Discrete Physics-Informed Neural Networks
by: Yin, Jichao, et al.
Published: (2025)
by: Yin, Jichao, et al.
Published: (2025)
Finding NeMo: Localizing Neurons Responsible For Memorization in Diffusion Models
by: Hintersdorf, Dominik, et al.
Published: (2024)
by: Hintersdorf, Dominik, et al.
Published: (2024)
C$^2$LEVA: Toward Comprehensive and Contamination-Free Language Model Evaluation
by: Li, Yanyang, et al.
Published: (2024)
by: Li, Yanyang, et al.
Published: (2024)
Online Video Understanding: OVBench and VideoChat-Online
by: Huang, Zhenpeng, et al.
Published: (2024)
by: Huang, Zhenpeng, et al.
Published: (2024)
Lying with Truths: Open-Channel Multi-Agent Collusion for Belief Manipulation via Generative Montage
by: Hu, Jinwei, et al.
Published: (2026)
by: Hu, Jinwei, et al.
Published: (2026)
Membership Inference on LLMs in the Wild
by: Yi, Jiatong, et al.
Published: (2026)
by: Yi, Jiatong, et al.
Published: (2026)
NeMo-map: Neural Implicit Flow Fields for Spatio-Temporal Motion Mapping
by: Zhu, Yufei, et al.
Published: (2025)
by: Zhu, Yufei, et al.
Published: (2025)
Omni-Video: Democratizing Unified Video Understanding and Generation
by: Tan, Zhiyu, et al.
Published: (2025)
by: Tan, Zhiyu, et al.
Published: (2025)
Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
by: Xu, Zhiyang, et al.
Published: (2026)
by: Xu, Zhiyang, et al.
Published: (2026)
NeMo: A Neuron-Level Modularizing-While-Training Approach for Decomposing DNN Models
by: Bi, Xiaohan, et al.
Published: (2025)
by: Bi, Xiaohan, et al.
Published: (2025)
Sequential-NIAH: A Needle-In-A-Haystack Benchmark for Extracting Sequential Needles from Long Contexts
by: Yu, Yifei, et al.
Published: (2025)
by: Yu, Yifei, et al.
Published: (2025)
Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs
by: Zhao, Zijia, et al.
Published: (2024)
by: Zhao, Zijia, et al.
Published: (2024)
ControlNeXt: Powerful and Efficient Control for Image and Video Generation
by: Peng, Bohao, et al.
Published: (2024)
by: Peng, Bohao, et al.
Published: (2024)
MoNeRF: Deformable Neural Rendering for Talking Heads via Latent Motion Navigation
by: X. Li, et al.
Published: (2024)
by: X. Li, et al.
Published: (2024)
VideoARM: Agentic Reasoning over Hierarchical Memory for Long-Form Video Understanding
by: Yin, Yufei, et al.
Published: (2025)
by: Yin, Yufei, et al.
Published: (2025)
VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
by: Wang, Jiapeng, et al.
Published: (2024)
by: Wang, Jiapeng, et al.
Published: (2024)
FAIL: Flow Matching Adversarial Imitation Learning for Image Generation
by: Ma, Yeyao, et al.
Published: (2026)
by: Ma, Yeyao, et al.
Published: (2026)
GP-NeRF: Generalized Perception NeRF for Context-Aware 3D Scene Understanding
by: Li, Hao, et al.
Published: (2023)
by: Li, Hao, et al.
Published: (2023)
VideoPro: Adaptive Program Reasoning for Long Video Understanding
by: Li, Chenglin, et al.
Published: (2025)
by: Li, Chenglin, et al.
Published: (2025)
InsertNeRF: Instilling Generalizability into NeRF with HyperNet Modules
by: Bao, Yanqi, et al.
Published: (2023)
by: Bao, Yanqi, et al.
Published: (2023)
Deep Video Discovery: Agentic Search with Tool Use for Long-form Video Understanding
by: Zhang, Xiaoyi, et al.
Published: (2025)
by: Zhang, Xiaoyi, et al.
Published: (2025)
Enhancing Subsequent Video Retrieval via Vision-Language Models (VLMs)
by: Duan, Yicheng, et al.
Published: (2025)
by: Duan, Yicheng, et al.
Published: (2025)
Similar Items
-
Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry Priors
by: Zheng, Duo, et al.
Published: (2025) -
Efficient-VLN: A Training-Efficient Vision-Language Navigation Model
by: Zheng, Duo, et al.
Published: (2025) -
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
by: Zheng, Duo, et al.
Published: (2024) -
Training Video Foundation Models with NVIDIA NeMo
by: Patel, Zeeshan, et al.
Published: (2025) -
Fine-grained Spatiotemporal Grounding on Egocentric Videos
by: Liang, Shuo, et al.
Published: (2025)