Guardado en:
| Autores principales: | Salamatian, Ali, Fuller, Anthony, Sarkar, Pritam, Green, James R., Sigal, Leonid, Shelhamer, Evan |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2605.06809 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
por: Fuller, Anthony, et al.
Publicado: (2025)
por: Fuller, Anthony, et al.
Publicado: (2025)
Thicker and Quicker: A Jumbo Token for Fast Plain Vision Transformers
por: Fuller, Anthony, et al.
Publicado: (2025)
por: Fuller, Anthony, et al.
Publicado: (2025)
LookSharp: Attention Entropy Minimization for Test-Time Adaptation
por: Mali, Yash, et al.
Publicado: (2025)
por: Mali, Yash, et al.
Publicado: (2025)
A Closer Look at In-Distribution vs. Out-of-Distribution Accuracy for Open-Set Test-time Adaptation
por: Li, Zefeng, et al.
Publicado: (2026)
por: Li, Zefeng, et al.
Publicado: (2026)
Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
por: Lowe, Scott C., et al.
Publicado: (2026)
por: Lowe, Scott C., et al.
Publicado: (2026)
Galileo: Learning Global & Local Features of Many Remote Sensing Modalities
por: Tseng, Gabriel, et al.
Publicado: (2025)
por: Tseng, Gabriel, et al.
Publicado: (2025)
LookHere: Vision Transformers with Directed Attention Generalize and Extrapolate
por: Fuller, Anthony, et al.
Publicado: (2024)
por: Fuller, Anthony, et al.
Publicado: (2024)
Self-alignment of Large Video Language Models with Refined Regularized Preference Optimization
por: Sarkar, Pritam, et al.
Publicado: (2025)
por: Sarkar, Pritam, et al.
Publicado: (2025)
VCRBench: Exploring Long-form Causal Reasoning Capabilities of Large Video Language Models
por: Sarkar, Pritam, et al.
Publicado: (2025)
por: Sarkar, Pritam, et al.
Publicado: (2025)
ChartGaze: Enhancing Chart Understanding in LVLMs with Eye-Tracking Guided Attention Refinement
por: Salamatian, Ali, et al.
Publicado: (2025)
por: Salamatian, Ali, et al.
Publicado: (2025)
The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding
por: Luo, Jiayun, et al.
Publicado: (2024)
por: Luo, Jiayun, et al.
Publicado: (2024)
What and When to Look?: Temporal Span Proposal Network for Video Relation Detection
por: Woo, Sangmin, et al.
Publicado: (2021)
por: Woo, Sangmin, et al.
Publicado: (2021)
Implicit and Explicit Commonsense for Multi-sentence Video Captioning
por: Chou, Shih-Han, et al.
Publicado: (2023)
por: Chou, Shih-Han, et al.
Publicado: (2023)
LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models
por: Shen, Yuxiang, et al.
Publicado: (2026)
por: Shen, Yuxiang, et al.
Publicado: (2026)
No One Knows the State of the Art in Geospatial Foundation Models
por: Corley, Isaac, et al.
Publicado: (2026)
por: Corley, Isaac, et al.
Publicado: (2026)
When to Think and When to Look: Uncertainty-Guided Lookback
por: Bi, Jing, et al.
Publicado: (2025)
por: Bi, Jing, et al.
Publicado: (2025)
StreamReady: Learning What to Answer and When in Long Streaming Videos
por: Azad, Shehreen, et al.
Publicado: (2026)
por: Azad, Shehreen, et al.
Publicado: (2026)
Show Me When and Where: Towards Referring Video Object Segmentation in the Wild
por: Gao, Mingqi, et al.
Publicado: (2026)
por: Gao, Mingqi, et al.
Publicado: (2026)
Factorized Video Autoencoders for Efficient Generative Modelling
por: Suhail, Mohammed, et al.
Publicado: (2024)
por: Suhail, Mohammed, et al.
Publicado: (2024)
What Happens When: Learning Temporal Orders of Events in Videos
por: Ahn, Daechul, et al.
Publicado: (2025)
por: Ahn, Daechul, et al.
Publicado: (2025)
All in One: A Unified Synthetic Data Pipeline for Multimodal Video Understanding
por: Rahman, Tanzila, et al.
Publicado: (2026)
por: Rahman, Tanzila, et al.
Publicado: (2026)
When and Where do Events Switch in Multi-Event Video Generation?
por: Liao, Ruotong, et al.
Publicado: (2025)
por: Liao, Ruotong, et al.
Publicado: (2025)
When Dance Video Archives Challenge Computer Vision
por: Colantoni, Philippe, et al.
Publicado: (2025)
por: Colantoni, Philippe, et al.
Publicado: (2025)
Preventing Catastrophic Forgetting through Memory Networks in Continuous Detection
por: Bhatt, Gaurav, et al.
Publicado: (2024)
por: Bhatt, Gaurav, et al.
Publicado: (2024)
ProtoTTA: Prototype-Guided Test-Time Adaptation
por: Abootorabi, Mohammad Mahdi, et al.
Publicado: (2026)
por: Abootorabi, Mohammad Mahdi, et al.
Publicado: (2026)
TAM-VT: Transformation-Aware Multi-scale Video Transformer for Segmentation and Tracking
por: Goyal, Raghav, et al.
Publicado: (2023)
por: Goyal, Raghav, et al.
Publicado: (2023)
AI-Generated Images: What Humans and Machines See When They Look at the Same Image
por: Poletti, Silvia, et al.
Publicado: (2026)
por: Poletti, Silvia, et al.
Publicado: (2026)
Learning When to Look: A Disentangled Curriculum for Strategic Perception in Multimodal Reasoning
por: Yang, Siqi, et al.
Publicado: (2025)
por: Yang, Siqi, et al.
Publicado: (2025)
Learning What Matters: Prioritized Concept Learning via Relative Error-driven Sample Selection
por: Chandhok, Shivam, et al.
Publicado: (2025)
por: Chandhok, Shivam, et al.
Publicado: (2025)
Spotlight: Identifying and Localizing Video Generation Errors Using VLMs
por: Chinchure, Aditya, et al.
Publicado: (2025)
por: Chinchure, Aditya, et al.
Publicado: (2025)
CoRDS: Coreset-based Representative and Diverse Selection for Streaming Video Understanding
por: Mahdizadeh, Ailar, et al.
Publicado: (2026)
por: Mahdizadeh, Ailar, et al.
Publicado: (2026)
How Animals Dance (When You're Not Looking)
por: Wang, Xiaojuan, et al.
Publicado: (2025)
por: Wang, Xiaojuan, et al.
Publicado: (2025)
When and What: Diffusion-Grounded VideoLLM with Entity Aware Segmentation for Long Video Understanding
por: Fang, Pengcheng, et al.
Publicado: (2025)
por: Fang, Pengcheng, et al.
Publicado: (2025)
When, Where, and What? A Novel Benchmark for Accident Anticipation and Localization with Large Language Models
por: Liao, Haicheng, et al.
Publicado: (2024)
por: Liao, Haicheng, et al.
Publicado: (2024)
Self-Soupervision: Cooking Model Soups without Labels
por: Fuller, Anthony, et al.
Publicado: (2026)
por: Fuller, Anthony, et al.
Publicado: (2026)
GUI Action Narrator: Where and When Did That Action Take Place?
por: Wu, Qinchen, et al.
Publicado: (2024)
por: Wu, Qinchen, et al.
Publicado: (2024)
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
por: Duan, Yuxiang, et al.
Publicado: (2025)
por: Duan, Yuxiang, et al.
Publicado: (2025)
SPIKE-RL: Video-LLMs meet Bayesian Surprise
por: Ravi, Sahithya, et al.
Publicado: (2025)
por: Ravi, Sahithya, et al.
Publicado: (2025)
ReservoirTTA: Prolonged Test-time Adaptation for Evolving and Recurring Domains
por: Vray, Guillaume, et al.
Publicado: (2025)
por: Vray, Guillaume, et al.
Publicado: (2025)
Response Wide Shut: Surprising Observations in Basic Vision Language Model Capabilities
por: Chandhok, Shivam, et al.
Publicado: (2024)
por: Chandhok, Shivam, et al.
Publicado: (2024)
Ejemplares similares
-
LookWhere? Efficient Visual Recognition by Learning Where to Look and What to See from Self-Supervision
por: Fuller, Anthony, et al.
Publicado: (2025) -
Thicker and Quicker: A Jumbo Token for Fast Plain Vision Transformers
por: Fuller, Anthony, et al.
Publicado: (2025) -
LookSharp: Attention Entropy Minimization for Test-Time Adaptation
por: Mali, Yash, et al.
Publicado: (2025) -
A Closer Look at In-Distribution vs. Out-of-Distribution Accuracy for Open-Set Test-time Adaptation
por: Li, Zefeng, et al.
Publicado: (2026) -
Self-Distillation of Hidden Layers for Self-Supervised Representation Learning
por: Lowe, Scott C., et al.
Publicado: (2026)