Gespeichert in:
| Hauptverfasser: | Li, Zhiqi, Chen, Guo, Liu, Shilong, Wang, Shihao, VS, Vibashan, Ji, Yishen, Lan, Shiyi, Zhang, Hao, Zhao, Yilin, Radhakrishnan, Subhashree, Chang, Nadine, Sapra, Karan, Deshmukh, Amala Sanjay, Rintamaki, Tuomas, Le, Matthieu, Karmanov, Ilia, Voegtle, Lukas, Fischer, Philipp, Huang, De-An, Roman, Timo, Lu, Tong, Alvarez, Jose M., Catanzaro, Bryan, Kautz, Jan, Tao, Andrew, Liu, Guilin, Yu, Zhiding |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2501.14818 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
von: Shi, Min, et al.
Veröffentlicht: (2024)
von: Shi, Min, et al.
Veröffentlicht: (2024)
Éclair -- Extracting Content and Layout with Integrated Reading Order for Documents
von: Karmanov, Ilia, et al.
Veröffentlicht: (2025)
von: Karmanov, Ilia, et al.
Veröffentlicht: (2025)
Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
von: Chen, Guo, et al.
Veröffentlicht: (2025)
von: Chen, Guo, et al.
Veröffentlicht: (2025)
Stateful Token Reduction for Long-Video Hybrid VLMs
von: Jiang, Jindong, et al.
Veröffentlicht: (2026)
von: Jiang, Jindong, et al.
Veröffentlicht: (2026)
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
von: Huang, De-An, et al.
Veröffentlicht: (2025)
von: Huang, De-An, et al.
Veröffentlicht: (2025)
LITA: Language Instructed Temporal-Localization Assistant
von: Huang, De-An, et al.
Veröffentlicht: (2024)
von: Huang, De-An, et al.
Veröffentlicht: (2024)
Hydra-MDP: End-to-end Multimodal Planning with Multi-target Hydra-Distillation
von: Li, Zhenxin, et al.
Veröffentlicht: (2024)
von: Li, Zhenxin, et al.
Veröffentlicht: (2024)
FaceXBench: Evaluating Multimodal LLMs on Face Understanding
von: Narayan, Kartik, et al.
Veröffentlicht: (2025)
von: Narayan, Kartik, et al.
Veröffentlicht: (2025)
Certainty and Uncertainty Guided Active Domain Adaptation
von: Safaei, Bardia, et al.
Veröffentlicht: (2025)
von: Safaei, Bardia, et al.
Veröffentlicht: (2025)
SegFace: Face Segmentation of Long-Tail Classes
von: Narayan, Kartik, et al.
Veröffentlicht: (2024)
von: Narayan, Kartik, et al.
Veröffentlicht: (2024)
AIDE: Agentically Improve Visual Language Model with Domain Experts
von: Chiu, Ming-Chang, et al.
Veröffentlicht: (2025)
von: Chiu, Ming-Chang, et al.
Veröffentlicht: (2025)
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding
von: Wang, Shihao, et al.
Veröffentlicht: (2026)
von: Wang, Shihao, et al.
Veröffentlicht: (2026)
What is Point Supervision Worth in Video Instance Segmentation?
von: Huang, Shuaiyi, et al.
Veröffentlicht: (2024)
von: Huang, Shuaiyi, et al.
Veröffentlicht: (2024)
Is Ego Status All You Need for Open-Loop End-to-End Autonomous Driving?
von: Li, Zhiqi, et al.
Veröffentlicht: (2023)
von: Li, Zhiqi, et al.
Veröffentlicht: (2023)
FaceXFormer: A Unified Transformer for Facial Analysis
von: Narayan, Kartik, et al.
Veröffentlicht: (2024)
von: Narayan, Kartik, et al.
Veröffentlicht: (2024)
Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
von: Zhang, Shaokun, et al.
Veröffentlicht: (2025)
von: Zhang, Shaokun, et al.
Veröffentlicht: (2025)
OMCAT: Omni Context Aware Transformer
von: Goel, Arushi, et al.
Veröffentlicht: (2024)
von: Goel, Arushi, et al.
Veröffentlicht: (2024)
PhyCritic: Multimodal Critic Models for Physical AI
von: Xiong, Tianyi, et al.
Veröffentlicht: (2026)
von: Xiong, Tianyi, et al.
Veröffentlicht: (2026)
StreamChat: Chatting with Streaming Video
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
von: Liu, Jihao, et al.
Veröffentlicht: (2024)
ImagineMap: Enhanced HD Map Construction with SD Maps
von: Ji, Yishen, et al.
Veröffentlicht: (2024)
von: Ji, Yishen, et al.
Veröffentlicht: (2024)
VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought
von: Man, Yunze, et al.
Veröffentlicht: (2025)
von: Man, Yunze, et al.
Veröffentlicht: (2025)
Zero-Shot Scene Understanding for Automatic Target Recognition Using Large Vision-Language Models
von: Ranasinghe, Yasiru, et al.
Veröffentlicht: (2025)
von: Ranasinghe, Yasiru, et al.
Veröffentlicht: (2025)
Mensch - Maske - Tier. Zu den Entstehungsbedingungen der Karikatur
von: Simone Voegtle
Veröffentlicht: (2017)
von: Simone Voegtle
Veröffentlicht: (2017)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
von: Wang, Shihao, et al.
Veröffentlicht: (2025)
OmniDrive: A Holistic Vision-Language Dataset for Autonomous Driving with Counterfactual Reasoning
von: Wang, Shihao, et al.
Veröffentlicht: (2024)
von: Wang, Shihao, et al.
Veröffentlicht: (2024)
NVIDIA Nemotron Parse 1.1
von: Chumachenko, Kateryna, et al.
Veröffentlicht: (2025)
von: Chumachenko, Kateryna, et al.
Veröffentlicht: (2025)
PARC: A Quantitative Framework Uncovering the Symmetries within Vision Language Models
von: Schmalfuss, Jenny, et al.
Veröffentlicht: (2025)
von: Schmalfuss, Jenny, et al.
Veröffentlicht: (2025)
Towards Multimodal Lifelong Understanding: A Dataset and Agentic Baseline
von: Chen, Guo, et al.
Veröffentlicht: (2026)
von: Chen, Guo, et al.
Veröffentlicht: (2026)
Hydra-NeXt: Robust Closed-Loop Driving with Open-Loop Training
von: Li, Zhenxin, et al.
Veröffentlicht: (2025)
von: Li, Zhenxin, et al.
Veröffentlicht: (2025)
NVLM: Open Frontier-Class Multimodal LLMs
von: Dai, Wenliang, et al.
Veröffentlicht: (2024)
von: Dai, Wenliang, et al.
Veröffentlicht: (2024)
PosSAM: Panoptic Open-vocabulary Segment Anything
von: VS, Vibashan, et al.
Veröffentlicht: (2024)
von: VS, Vibashan, et al.
Veröffentlicht: (2024)
LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
von: Man, Yunze, et al.
Veröffentlicht: (2025)
von: Man, Yunze, et al.
Veröffentlicht: (2025)
3D-Layout-R1: Structured Reasoning for Language-Instructed Spatial Editing
von: Zhen, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhen, Haoyu, et al.
Veröffentlicht: (2026)
StereoDETR: Stereo-based Transformer for 3D Object Detection
von: Mu, Shiyi, et al.
Veröffentlicht: (2025)
von: Mu, Shiyi, et al.
Veröffentlicht: (2025)
Old age, high risk medication, polypharmacy: a ‘trilogy’ of risks in older patients with atrial fibrillation
von: Yishen WANG
Veröffentlicht: (2016)
von: Yishen WANG
Veröffentlicht: (2016)
Image-to-Text for Medical Reports Using Adaptive Co-Attention and Triple-LSTM Module
von: Liu, Yishen
Veröffentlicht: (2025)
von: Liu, Yishen
Veröffentlicht: (2025)
Developing an Ontology for AI Act Fundamental Rights Impact Assessments
von: Rintamaki, Tytti, et al.
Veröffentlicht: (2024)
von: Rintamaki, Tytti, et al.
Veröffentlicht: (2024)
Towards An Automated AI Act FRIA Tool That Can Reuse GDPR's DPIA
von: Rintamaki, Tytti, et al.
Veröffentlicht: (2024)
von: Rintamaki, Tytti, et al.
Veröffentlicht: (2024)
Slow-Fast Architecture for Video Multi-Modal Large Language Models
von: Shi, Min, et al.
Veröffentlicht: (2025)
von: Shi, Min, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
von: Shi, Min, et al.
Veröffentlicht: (2024) -
Éclair -- Extracting Content and Layout with Integrated Reading Order for Documents
von: Karmanov, Ilia, et al.
Veröffentlicht: (2025) -
Eagle 2.5: Boosting Long-Context Post-Training for Frontier Vision-Language Models
von: Chen, Guo, et al.
Veröffentlicht: (2025) -
Stateful Token Reduction for Long-Video Hybrid VLMs
von: Jiang, Jindong, et al.
Veröffentlicht: (2026) -
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
von: Huang, De-An, et al.
Veröffentlicht: (2025)