PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Cho, Jang Hyun, Madotto, Andrea, Mavroudi, Effrosyni, Afouras, Triantafyllos, Nagarajan, Tushar, Maaz, Muhammad, Song, Yale, Ma, Tengyu, Hu, Shuming, Jain, Suyog, Martin, Miguel, Wang, Huiyu, Rasheed, Hanoona, Sun, Peize, Huang, Po-Yao, Bolya, Daniel, Ravi, Nikhila, Jain, Shashank, Stark, Tammy, Moon, Shane, Damavandi, Babak, Lee, Vivian, Westbury, Andrew, Khan, Salman, Krähenbühl, Philipp, Dollár, Piotr, Torresani, Lorenzo, Grauman, Kristen, Feichtenhofer, Christoph |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
by: Pramanick, Shraman, et al.
Published: (2025)
by: Pramanick, Shraman, et al.
Published: (2025)
Perception Encoder: The best visual embeddings are not at the output of the network
by: Bolya, Daniel, et al.
Published: (2025)
by: Bolya, Daniel, et al.
Published: (2025)
ViewBridge: Curriculum Knowledge Distillation for Activity View-Invariance Under Extreme Viewpoint Changes
by: Somayazulu, Arjun, et al.
Published: (2025)
by: Somayazulu, Arjun, et al.
Published: (2025)
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
by: Maaz, Muhammad, et al.
Published: (2024)
by: Maaz, Muhammad, et al.
Published: (2024)
VoiceVector: Multimodal Enrolment Vectors for Speaker Separation
by: Rahimi, Akam, et al.
Published: (2025)
by: Rahimi, Akam, et al.
Published: (2025)
Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation
by: Rahimi, Akam, et al.
Published: (2025)
by: Rahimi, Akam, et al.
Published: (2025)
Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
by: Maaz, Muhammad, et al.
Published: (2023)
by: Maaz, Muhammad, et al.
Published: (2023)
Video-R2: Reinforcing Consistent and Grounded Reasoning in Multimodal Language Models
by: Maaz, Muhammad, et al.
Published: (2025)
by: Maaz, Muhammad, et al.
Published: (2025)
Collecting Consistently High Quality Object Tracks with Minimal Human Involvement by Using Self-Supervised Learning to Detect Tracker Errors
by: Anjum, Samreen, et al.
Published: (2024)
by: Anjum, Samreen, et al.
Published: (2024)
Counterfactual Multi-Agent Policy Gradients
by: Foerster, Jakob, et al.
Published: (2017)
by: Foerster, Jakob, et al.
Published: (2017)
UniversalVTG: A Universal and Lightweight Foundation Model for Video Temporal Grounding
by: An, Joungbin, et al.
Published: (2026)
by: An, Joungbin, et al.
Published: (2026)
Video-CoM: Interactive Video Reasoning via Chain of Manipulations
by: Rasheed, Hanoona, et al.
Published: (2025)
by: Rasheed, Hanoona, et al.
Published: (2025)
UNETR++: Delving into Efficient and Accurate 3D Medical Image Segmentation
by: Shaker, Abdelrahman, et al.
Published: (2022)
by: Shaker, Abdelrahman, et al.
Published: (2022)
SAM 3: Segment Anything with Concepts
by: Carion, Nicolas, et al.
Published: (2025)
by: Carion, Nicolas, et al.
Published: (2025)
VideoMathQA: Benchmarking Mathematical Reasoning via Multimodal Understanding in Videos
by: Rasheed, Hanoona, et al.
Published: (2025)
by: Rasheed, Hanoona, et al.
Published: (2025)
VITED: Video Temporal Evidence Distillation
by: Lu, Yujie, et al.
Published: (2025)
by: Lu, Yujie, et al.
Published: (2025)
PALO: A Polyglot Large Multimodal Model for 5B People
by: Maaz, Muhammad, et al.
Published: (2024)
by: Maaz, Muhammad, et al.
Published: (2024)
Planting arrangement, nitrogen resources and plant density on some vegetative characteristics of Melissa officinalis
by: Zahra Damavandi
Published: (2015)
by: Zahra Damavandi
Published: (2015)
Classification of invariant tight contact structures on the 3-space, -ball and -sphere
by: Torresani, Mirko
Published: (2026)
by: Torresani, Mirko
Published: (2026)
The practicality of pure reason. A normative defence of Kant's theory of moral motivation
by: Triantafyllos Gkouvas
Published: (2011)
by: Triantafyllos Gkouvas
Published: (2011)
Proactive Assistant Dialogue Generation from Streaming Egocentric Videos
by: Zhang, Yichi, et al.
Published: (2025)
by: Zhang, Yichi, et al.
Published: (2025)
GEOMETRIA ÓSSEA E ATIVIDADE FÍSICA EM CRIANÇAS E ADOLESCENTES: REVISÃO SISTEMÁTICA
by: Tathyane Krahenbühl
Published: (2018)
by: Tathyane Krahenbühl
Published: (2018)
The use of the additional field player in handball: analysis of the Rio 2016 Olympic Games
by: Tathyane Krahenbühl
Published: (2019)
by: Tathyane Krahenbühl
Published: (2019)
Fatores que influenciam a massa óssea de crianças e adolescentes saudáveis mensurada pelo ultrassom quantitativo de falanges: revisão sistemática
by: Tathyane Krahenbühl
Published: (2014)
by: Tathyane Krahenbühl
Published: (2014)
SAM 2: Segment Anything in Images and Videos
by: Ravi, Nikhila, et al.
Published: (2024)
by: Ravi, Nikhila, et al.
Published: (2024)
GLaMM: Pixel Grounding Large Multimodal Model
by: Rasheed, Hanoona, et al.
Published: (2023)
by: Rasheed, Hanoona, et al.
Published: (2023)
ESL COMMUNITY PARTNER MEMBERS' PERCEPTIONS OF HIGHER EDUCATION SERVICE-LEARNING: A QUALITATIVE STUDY
by: Dollar, Natalya
Published: (2022)
by: Dollar, Natalya
Published: (2022)
Asian century of multi-polar century? / David Dollar
by: Dollar, David
Published: (2007)
by: Dollar, David
Published: (2007)
Sowing and reaping : institutional quality and project outcomes in developing countries / David Dollar, Victoria Levin
by: Dollar, David
Published: (2005)
by: Dollar, David
Published: (2005)
The increasing selectivity of foreign aid, 1984-2002 / David Dollar, Victoria Levin
by: Dollar, David
Published: (2004)
by: Dollar, David
Published: (2004)
Growth is good for the poor / David Dollar, Aart Kraay
by: Dollar, David
Published: (2001)
by: Dollar, David
Published: (2001)
Trade, growth, and poverty / David Dollar, Aart Kraay
by: Dollar, David
Published: (2001)
by: Dollar, David
Published: (2001)
The search for the key : id, investment and policies in Africa / David Dollar, William Easterly
by: Dollar, David
Published: (1999)
by: Dollar, David
Published: (1999)
Institutions, trade, and growth : revisiting the evidence / David Dollar, Aart Kraay
by: Dollar, David
Published: (2003)
by: Dollar, David
Published: (2003)
Globalization, poverty, and inequality since 1980 / David Dollar
by: Dollar, David
Published: (2004)
by: Dollar, David
Published: (2004)
Neither a borrower nor a lender : does China's zero net foreign asset position make economic sense ? / David Dollar, Aart Kraay
by: Dollar, David
Published: (2005)
by: Dollar, David
Published: (2005)
Conceptualizing grandiose and vulnerable narcissism as alternative status‐seeking strategies: Insights from hierometer theory
by: Nikhila Mahadevan
Published: (2024)
by: Nikhila Mahadevan
Published: (2024)
A Constant Error, Revisited: A New Explanation of the Halo Effect
by: Chris Westbury, et al.
Published: (2024)
by: Chris Westbury, et al.
Published: (2024)
Blaschke operations on log-concave functions and affine isoperimetric inequalities
by: Chasioti, Effrosyni, et al.
Published: (2026)
by: Chasioti, Effrosyni, et al.
Published: (2026)
Step Differences in Instructional Video
by: Nagarajan, Tushar, et al.
Published: (2024)
by: Nagarajan, Tushar, et al.
Published: (2024)
Similar Items
-
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
by: Pramanick, Shraman, et al.
Published: (2025) -
Perception Encoder: The best visual embeddings are not at the output of the network
by: Bolya, Daniel, et al.
Published: (2025) -
ViewBridge: Curriculum Knowledge Distillation for Activity View-Invariance Under Extreme Viewpoint Changes
by: Somayazulu, Arjun, et al.
Published: (2025) -
VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
by: Maaz, Muhammad, et al.
Published: (2024) -
VoiceVector: Multimodal Enrolment Vectors for Speaker Separation
by: Rahimi, Akam, et al.
Published: (2025)