Perception First: A Frontier Native-Video Model with Self-Consistency for Implicit Video Question Answering
Fuente:
arXiv
Guardado en:
| Autor principal: | Alavi, Ali |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
TLG: Temporal-Logic Grounding for Video Question Answering via Source-Annotation Reconstruction and Category-Targeted Reasoning
por: Alavi, Ali
Publicado: (2026)
por: Alavi, Ali
Publicado: (2026)
Reason, Retrieve, Re-rank: A Zero-Shot Reasoning-Aware Framework for Composed Video Retrieval
por: Alavi, Ali
Publicado: (2026)
por: Alavi, Ali
Publicado: (2026)
Improving Video Question Answering through query-based frame selection
por: Patil, Himanshu, et al.
Publicado: (2026)
por: Patil, Himanshu, et al.
Publicado: (2026)
Towards Fine-Grained Video Question Answering
por: Dai, Wei, et al.
Publicado: (2025)
por: Dai, Wei, et al.
Publicado: (2025)
CinePile: A Long Video Question Answering Dataset and Benchmark
por: Rawal, Ruchit, et al.
Publicado: (2024)
por: Rawal, Ruchit, et al.
Publicado: (2024)
Bridging Vision Language Models and Symbolic Grounding for Video Question Answering
por: Ma, Haodi, et al.
Publicado: (2025)
por: Ma, Haodi, et al.
Publicado: (2025)
Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions
por: Indrehus, Kjetil, et al.
Publicado: (2026)
por: Indrehus, Kjetil, et al.
Publicado: (2026)
MoReVQA: Exploring Modular Reasoning Models for Video Question Answering
por: Min, Juhong, et al.
Publicado: (2024)
por: Min, Juhong, et al.
Publicado: (2024)
Advancing Egocentric Video Question Answering with Multimodal Large Language Models
por: Patel, Alkesh, et al.
Publicado: (2025)
por: Patel, Alkesh, et al.
Publicado: (2025)
Semantically Consistent Video Inpainting with Conditional Diffusion Models
por: Green, Dylan, et al.
Publicado: (2024)
por: Green, Dylan, et al.
Publicado: (2024)
Questioning the Stability of Visual Question Answering
por: Rosenfeld, Amir, et al.
Publicado: (2025)
por: Rosenfeld, Amir, et al.
Publicado: (2025)
An Image Grid Can Be Worth a Video: Zero-shot Video Question Answering Using a VLM
por: Kim, Wonkyun, et al.
Publicado: (2024)
por: Kim, Wonkyun, et al.
Publicado: (2024)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
por: Shang, Chuyi, et al.
Publicado: (2024)
por: Shang, Chuyi, et al.
Publicado: (2024)
VideoNSA: Native Sparse Attention Scales Video Understanding
por: Song, Enxin, et al.
Publicado: (2025)
por: Song, Enxin, et al.
Publicado: (2025)
TPCL: Task Progressive Curriculum Learning for Robust Visual Question Answering
por: Akl, Ahmed, et al.
Publicado: (2024)
por: Akl, Ahmed, et al.
Publicado: (2024)
DocVXQA: Context-Aware Visual Explanations for Document Question Answering
por: Souibgui, Mohamed Ali, et al.
Publicado: (2025)
por: Souibgui, Mohamed Ali, et al.
Publicado: (2025)
WorldPack: Compressed Memory Improves Spatial Consistency in Video World Modeling
por: Oshima, Yuta, et al.
Publicado: (2025)
por: Oshima, Yuta, et al.
Publicado: (2025)
Self-Refining Video Sampling
por: Jang, Sangwon, et al.
Publicado: (2026)
por: Jang, Sangwon, et al.
Publicado: (2026)
STLLaVA-Med: Self-Training Large Language and Vision Assistant for Medical Question-Answering
por: Sun, Guohao, et al.
Publicado: (2024)
por: Sun, Guohao, et al.
Publicado: (2024)
Vidu: a Highly Consistent, Dynamic and Skilled Text-to-Video Generator with Diffusion Models
por: Bao, Fan, et al.
Publicado: (2024)
por: Bao, Fan, et al.
Publicado: (2024)
Describe Anything Model for Visual Question Answering on Text-rich Images
por: Vu, Yen-Linh, et al.
Publicado: (2025)
por: Vu, Yen-Linh, et al.
Publicado: (2025)
RSAdapter: Adapting Multimodal Models for Remote Sensing Visual Question Answering
por: Wang, Yuduo, et al.
Publicado: (2023)
por: Wang, Yuduo, et al.
Publicado: (2023)
RAVEN: Query-Guided Representation Alignment for Question Answering over Audio, Video, Embedded Sensors, and Natural Language
por: Biswas, Subrata, et al.
Publicado: (2025)
por: Biswas, Subrata, et al.
Publicado: (2025)
BERT-VQA: Visual Question Answering on Plots
por: Vu, Tai, et al.
Publicado: (2025)
por: Vu, Tai, et al.
Publicado: (2025)
LifelongMemory: Leveraging LLMs for Answering Queries in Long-form Egocentric Videos
por: Wang, Ying, et al.
Publicado: (2023)
por: Wang, Ying, et al.
Publicado: (2023)
Vision-Language Models for Medical Report Generation and Visual Question Answering: A Review
por: Hartsock, Iryna, et al.
Publicado: (2024)
por: Hartsock, Iryna, et al.
Publicado: (2024)
Sparse Forcing: Native Trainable Sparse Attention for Real-time Autoregressive Diffusion Video Generation
por: Xu, Boxun, et al.
Publicado: (2026)
por: Xu, Boxun, et al.
Publicado: (2026)
CETCAM: Camera-Controllable Video Generation via Consistent and Extensible Tokenization
por: Zhao, Zelin, et al.
Publicado: (2025)
por: Zhao, Zelin, et al.
Publicado: (2025)
SViQA: A Unified Speech-Vision Multimodal Model for Textless Visual Question Answering
por: Li, Bingxin
Publicado: (2025)
por: Li, Bingxin
Publicado: (2025)
Physics-guided Shape-from-Template: Monocular Video Perception through Neural Surrogate Models
por: Stotko, David, et al.
Publicado: (2023)
por: Stotko, David, et al.
Publicado: (2023)
VideoGPA: Distilling Geometry Priors for 3D-Consistent Video Generation
por: Du, Hongyang, et al.
Publicado: (2026)
por: Du, Hongyang, et al.
Publicado: (2026)
T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation
por: Guo, Xuyang, et al.
Publicado: (2025)
por: Guo, Xuyang, et al.
Publicado: (2025)
First Place Solution to the Multiple-choice Video QA Track of The Second Perception Test Challenge
por: Peng, Yingzhe, et al.
Publicado: (2024)
por: Peng, Yingzhe, et al.
Publicado: (2024)
COREVQA: A Crowd Observation and Reasoning Entailment Visual Question Answering Benchmark
por: Chintapatla, Ishant, et al.
Publicado: (2025)
por: Chintapatla, Ishant, et al.
Publicado: (2025)
Exploring the Effectiveness of Object-Centric Representations in Visual Question Answering: Comparative Insights with Foundation Models
por: Mamaghan, Amir Mohammad Karimi, et al.
Publicado: (2024)
por: Mamaghan, Amir Mohammad Karimi, et al.
Publicado: (2024)
Cross-modal Causal Relation Alignment for Video Question Grounding
por: Chen, Weixing, et al.
Publicado: (2025)
por: Chen, Weixing, et al.
Publicado: (2025)
Graph2Video: Leveraging Video Models to Model Dynamic Graph Evolution
por: Liu, Hua, et al.
Publicado: (2026)
por: Liu, Hua, et al.
Publicado: (2026)
Privacy-Aware Document Visual Question Answering
por: Tito, Rubèn, et al.
Publicado: (2023)
por: Tito, Rubèn, et al.
Publicado: (2023)
Video Diffusion Models: A Survey
por: Melnik, Andrew, et al.
Publicado: (2024)
por: Melnik, Andrew, et al.
Publicado: (2024)
VideoSAM: A Large Vision Foundation Model for High-Speed Video Segmentation
por: Maduabuchi, Chika, et al.
Publicado: (2024)
por: Maduabuchi, Chika, et al.
Publicado: (2024)
Ejemplares similares
-
TLG: Temporal-Logic Grounding for Video Question Answering via Source-Annotation Reconstruction and Category-Targeted Reasoning
por: Alavi, Ali
Publicado: (2026) -
Reason, Retrieve, Re-rank: A Zero-Shot Reasoning-Aware Framework for Composed Video Retrieval
por: Alavi, Ali
Publicado: (2026) -
Improving Video Question Answering through query-based frame selection
por: Patil, Himanshu, et al.
Publicado: (2026) -
Towards Fine-Grained Video Question Answering
por: Dai, Wei, et al.
Publicado: (2025) -
CinePile: A Long Video Question Answering Dataset and Benchmark
por: Rawal, Ruchit, et al.
Publicado: (2024)