See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Sun, Boyuan, Yin, Bowen, Li, Yuanming, Wei, Xihan, Hou, Qibin |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
by: Sun, Boyuan, et al.
Published: (2025)
by: Sun, Boyuan, et al.
Published: (2025)
See What I Mean? Mobile Eye-Perspective Rendering for Optical See-through Head-mounted Displays
by: Emsenhuber, Gerlinde, et al.
Published: (2025)
by: Emsenhuber, Gerlinde, et al.
Published: (2025)
See What I Mean? Expressiveness and Clarity in Robot Display Design
by: Ebisu, Matthew, et al.
Published: (2025)
by: Ebisu, Matthew, et al.
Published: (2025)
"See What I Imagine, Imagine What I See": Human-AI Co-Creation System for 360$^\circ$ Panoramic Video Generation in VR
by: Wen, Yunge
Published: (2025)
by: Wen, Yunge
Published: (2025)
See What I See: An Attention-Guiding eHMI Approach for Autonomous Vehicles
by: Li, Jialong, et al.
Published: (2026)
by: Li, Jialong, et al.
Published: (2026)
Do Models See in Line with Human Vision? Probing the Correspondence Between LVLM Representations and EEG Signals
by: Xiao, Xin, et al.
Published: (2026)
by: Xiao, Xin, et al.
Published: (2026)
Do You See What I See? A Qualitative Study Eliciting High-Level Visualization Comprehension
by: Quadri, Ghulam Jilani, et al.
Published: (2024)
by: Quadri, Ghulam Jilani, et al.
Published: (2024)
Accessible Fine-grained Data Representation via Spatial Audio
by: Liu, Can, et al.
Published: (2026)
by: Liu, Can, et al.
Published: (2026)
I See You: Teacher Analytics with GPT-4 Vision-Powered Observational Assessment
by: Lee, Unggi, et al.
Published: (2024)
by: Lee, Unggi, et al.
Published: (2024)
"What I Sign Is Not What I See": Towards Explainable and Trustworthy Cryptocurrency Wallet Signatures
by: Qin, Yuyang, et al.
Published: (2026)
by: Qin, Yuyang, et al.
Published: (2026)
Do Vision-Language Models See Visualizations Like Humans? Alignment in Chart Categorization
by: Gyarmati, Péter Ferenc, et al.
Published: (2025)
by: Gyarmati, Péter Ferenc, et al.
Published: (2025)
V-CASS: Vision-context-aware Expressive Speech Synthesis for Enhancing User Understanding of Videos
by: Wang, Qixin, et al.
Published: (2025)
by: Wang, Qixin, et al.
Published: (2025)
"I Don't Have Faith in the Developers to Use My Feedback": Understanding Player Values and Expectancy for Reporting Systems in Video Games
by: Yin, Michael, et al.
Published: (2026)
by: Yin, Michael, et al.
Published: (2026)
Believing without Seeing: Quality Scores for Contextualizing Vision-Language Model Explanations
by: He, Keyu, et al.
Published: (2025)
by: He, Keyu, et al.
Published: (2025)
What Sensors See, What People Feel: An Exploratory Study of Subjective Collaboration Perception in Mixed Reality
by: Chandio, Yasra, et al.
Published: (2025)
by: Chandio, Yasra, et al.
Published: (2025)
Uncovering the EEG Temporal Representation of Low-dimensional Object Properties
by: Tang, Jiahua, et al.
Published: (2025)
by: Tang, Jiahua, et al.
Published: (2025)
What Does a Meow Mean? In Search of Intuitively Understandable Communication by a Nonverbal Companion Robot
by: Chi, Vivienne Bihe, et al.
Published: (2026)
by: Chi, Vivienne Bihe, et al.
Published: (2026)
Understanding Pedestrian Gesture Misrecognition: Insights from Vision-Language Model Reasoning
by: Tran, Tram Thi Minh, et al.
Published: (2025)
by: Tran, Tram Thi Minh, et al.
Published: (2025)
Understanding the Impact of Referent Design on Scale Perception in Immersive Data Visualization
by: Hou, Yihan, et al.
Published: (2024)
by: Hou, Yihan, et al.
Published: (2024)
The Neural-Wave Quick Escape Manual 2036: A Field Guide to Adversarial Living in the Era of "Empathic" AIoT
by: Gu, Boyuan, et al.
Published: (2026)
by: Gu, Boyuan, et al.
Published: (2026)
"It's Kind of Context Dependent": Understanding Blind and Low Vision People's Video Accessibility Preferences Across Viewing Scenarios
by: Jiang, Lucy, et al.
Published: (2024)
by: Jiang, Lucy, et al.
Published: (2024)
I Was Blind but Now I See: Implementing Vision-Enabled Dialogue in Social Robots
by: Abbo, Giulio Antonio, et al.
Published: (2023)
by: Abbo, Giulio Antonio, et al.
Published: (2023)
"If I were in Space": Understanding and Adapting to Social Isolation through Designing Collaborative Narratives
by: Gong, Qi, et al.
Published: (2025)
by: Gong, Qi, et al.
Published: (2025)
Align-to-Scale: Mode Switching Technique for Unimanual 3D Object Manipulation with Gaze-Hand-Object Alignment in Extended Reality
by: Kim, Min-yung, et al.
Published: (2026)
by: Kim, Min-yung, et al.
Published: (2026)
Seeing Twice: How Side-by-Side T2I Comparison Changes Auditing Strategies
by: Maldaner, Matheus Kunzler, et al.
Published: (2025)
by: Maldaner, Matheus Kunzler, et al.
Published: (2025)
The "Huh?" Button: Improving Understanding in Educational Videos with Large Language Models
by: Ruf, Boris, et al.
Published: (2024)
by: Ruf, Boris, et al.
Published: (2024)
What Do We Mean When We Talk About Data Storytelling?
by: Yang, Leni, et al.
Published: (2025)
by: Yang, Leni, et al.
Published: (2025)
What If Moderation Didn't Mean Suppression? A Case for Personalized Content Transformation
by: Rashed, Rayhan, et al.
Published: (2025)
by: Rashed, Rayhan, et al.
Published: (2025)
"Mapping What I Feel": Understanding Affective Geovisualization Design Through the Lens of People-Place Relationships
by: Lan, Xingyu, et al.
Published: (2025)
by: Lan, Xingyu, et al.
Published: (2025)
Evaluating Perceptual Deviations in Video See-Through Head-Mounted Displays while Utilizing Physical Touchscreens
by: de Lange, Rudy De-Xin, et al.
Published: (2024)
by: de Lange, Rudy De-Xin, et al.
Published: (2024)
What People See (and Miss) About Generative AI Risks: Perceptions of Failures, Risks, and Who Should Address Them
by: Li, Megan, et al.
Published: (2026)
by: Li, Megan, et al.
Published: (2026)
Seeing the Reasoning: How LLM Rationales Influence User Trust and Decision-Making in Factual Verification Tasks
by: Sun, Xin, et al.
Published: (2026)
by: Sun, Xin, et al.
Published: (2026)
Alexa, I Wanna See You: Envisioning Smart Home Assistants for the Deaf and Hard-of-Hearing
by: Maria, Tyrone Justin Sta., et al.
Published: (2024)
by: Maria, Tyrone Justin Sta., et al.
Published: (2024)
Aligning Language Models with Demonstrated Feedback
by: Shaikh, Omar, et al.
Published: (2024)
by: Shaikh, Omar, et al.
Published: (2024)
DesignMinds: Enhancing Video-Based Design Ideation with Vision-Language Model and Context-Injected Large Language Model
by: He, Tianhao, et al.
Published: (2024)
by: He, Tianhao, et al.
Published: (2024)
Understanding User Experience in Large Language Model Interactions
by: Wang, Jiayin, et al.
Published: (2024)
by: Wang, Jiayin, et al.
Published: (2024)
LLM4Brain: Training a Large Language Model for Brain Video Understanding
by: Zheng, Ruizhe, et al.
Published: (2024)
by: Zheng, Ruizhe, et al.
Published: (2024)
Looking Together $\neq$ Seeing the Same Thing: Understanding Surgeons' Visual Needs During Intra-operative Coordination and Instruction
by: Popov, Vitaliy, et al.
Published: (2024)
by: Popov, Vitaliy, et al.
Published: (2024)
WristSonic: Enabling Fine-grained Hand-Face Interactions on Smartwatches Using Active Acoustic Sensing
by: Mahmud, Saif, et al.
Published: (2024)
by: Mahmud, Saif, et al.
Published: (2024)
"What If Smart Homes Could See Our Homes?": Exploring DIY Smart Home Building Experiences with VLM-Based Camera Sensors
by: Yun, Sojeong, et al.
Published: (2025)
by: Yun, Sojeong, et al.
Published: (2025)
Similar Items
-
LLaVA-Scissor: Token Compression with Semantic Connected Components for Video LLMs
by: Sun, Boyuan, et al.
Published: (2025) -
See What I Mean? Mobile Eye-Perspective Rendering for Optical See-through Head-mounted Displays
by: Emsenhuber, Gerlinde, et al.
Published: (2025) -
See What I Mean? Expressiveness and Clarity in Robot Display Design
by: Ebisu, Matthew, et al.
Published: (2025) -
"See What I Imagine, Imagine What I See": Human-AI Co-Creation System for 360$^\circ$ Panoramic Video Generation in VR
by: Wen, Yunge
Published: (2025) -
See What I See: An Attention-Guiding eHMI Approach for Autonomous Vehicles
by: Li, Jialong, et al.
Published: (2026)