UOUO: Uncontextualized Uncommon Objects for Measuring Knowledge Horizons of Vision Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Pi, Xinyu, Wu, Mingyuan, Jiang, Jize, Zheng, Haozhen, Tian, Beitong, Zhai, Chengxiang, Nahrstedt, Klara, Hu, Zhiting |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning
by: Wu, Mingyuan, et al.
Published: (2025)
by: Wu, Mingyuan, et al.
Published: (2025)
Spatio-Temporal LLM: Reasoning about Environments and Actions
by: Zheng, Haozhen, et al.
Published: (2025)
by: Zheng, Haozhen, et al.
Published: (2025)
AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
by: Tian, Beitong, et al.
Published: (2025)
by: Tian, Beitong, et al.
Published: (2025)
AquaScope: Reliable Underwater Image Transmission on Mobile Devices
by: Tian, Beitong, et al.
Published: (2025)
by: Tian, Beitong, et al.
Published: (2025)
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
by: Wu, Mingyuan, et al.
Published: (2025)
by: Wu, Mingyuan, et al.
Published: (2025)
Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking
by: Yang, Jingcheng, et al.
Published: (2026)
by: Yang, Jingcheng, et al.
Published: (2026)
TraceNet: Segment one thing efficiently
by: Wu, Mingyuan, et al.
Published: (2024)
by: Wu, Mingyuan, et al.
Published: (2024)
FDM-Bench: A Comprehensive Benchmark for Evaluating Large Language Models in Additive Manufacturing Tasks
by: Eslaminia, Ahmadreza, et al.
Published: (2024)
by: Eslaminia, Ahmadreza, et al.
Published: (2024)
Aha Moment Revisited: Are VLMs Truly Capable of Self Verification in Inference-time Scaling?
by: Wu, Mingyuan, et al.
Published: (2025)
by: Wu, Mingyuan, et al.
Published: (2025)
QoS-QoE Translation with Large Language Model
by: Yu, Yingjie, et al.
Published: (2026)
by: Yu, Yingjie, et al.
Published: (2026)
ACT360: An Efficient 360-Degree Action Detection and Summarization Framework for Mission-Critical Training and Debriefing
by: Tiwari, Aditi, et al.
Published: (2025)
by: Tiwari, Aditi, et al.
Published: (2025)
Performance Characterization of Containers in Edge Computing
by: Gupta, Ragini, et al.
Published: (2025)
by: Gupta, Ragini, et al.
Published: (2025)
Spreadsheet-RL: Advancing Large Language Model Agents on Realistic Spreadsheet Tasks via Reinforcement Learning
by: Chi, Banghao, et al.
Published: (2026)
by: Chi, Banghao, et al.
Published: (2026)
Revisiting Disaggregated Large Language Model Serving for Performance and Energy Implications
by: Li, Jiaxi, et al.
Published: (2025)
by: Li, Jiaxi, et al.
Published: (2025)
Report on the NSF Workshop on Sustainable Computing for Sustainability (NSF WSCS 2024)
by: Guérin, Roch, et al.
Published: (2024)
by: Guérin, Roch, et al.
Published: (2024)
Pseudo Dataset Generation for Out-of-Domain Multi-Camera View Recommendation
by: Lee, Kuan-Ying, et al.
Published: (2024)
by: Lee, Kuan-Ying, et al.
Published: (2024)
Competence-Based Analysis of Language Models
by: Davies, Adam, et al.
Published: (2023)
by: Davies, Adam, et al.
Published: (2023)
EcoLens: Leveraging Multi-Objective Bayesian Optimization for Energy-Efficient Video Processing on Edge Devices
by: Civjan, Benjamin, et al.
Published: (2025)
by: Civjan, Benjamin, et al.
Published: (2025)
Federated Transfer Learning with Task Personalization for Condition Monitoring in Ultrasonic Metal Welding
by: Eslaminia, Ahmadreza, et al.
Published: (2024)
by: Eslaminia, Ahmadreza, et al.
Published: (2024)
Viewport-based Neural 360° Image Compression
by: Liao, Jingwei, et al.
Published: (2026)
by: Liao, Jingwei, et al.
Published: (2026)
Generating Uncontextualized and Contextualized Questions for Document-Level Event Argument Extraction
by: Uddin, Md Nayem, et al.
Published: (2024)
by: Uddin, Md Nayem, et al.
Published: (2024)
Adaptive Unknown Fault Detection and Few-Shot Continual Learning for Condition Monitoring in Ultrasonic Metal Welding
by: Eslaminia, Ahmadreza, et al.
Published: (2026)
by: Eslaminia, Ahmadreza, et al.
Published: (2026)
Competition between surficial and volumetric diffusion in sintering TiO 2 polymorphs by molecular dynamics simulation
by: Jiang Li, et al.
Published: (2024)
by: Jiang Li, et al.
Published: (2024)
Enhancing Neural Adaptive Wireless Video Streaming via Lower-Layer Information Exposure and Online Tuning
by: Zhao, Lingzhi, et al.
Published: (2025)
by: Zhao, Lingzhi, et al.
Published: (2025)
Towards General Continuous Memory for Vision-Language Models
by: Wu, Wenyi, et al.
Published: (2025)
by: Wu, Wenyi, et al.
Published: (2025)
UAV-Track VLA: Embodied Aerial Tracking via Vision-Language-Action Models
by: Zhang, Qiyao, et al.
Published: (2026)
by: Zhang, Qiyao, et al.
Published: (2026)
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation
by: Sun, Jianli, et al.
Published: (2026)
by: Sun, Jianli, et al.
Published: (2026)
“I'm Gonna Always Make Everything OK for Them”: Rehabilitative Veneer, Stability Maintenance, and Offenders' Perceptions of Procedural (In)Justice Within Chinese Community Corrections
by: Jize Jiang, et al.
Published: (2025)
by: Jize Jiang, et al.
Published: (2025)
Leveraging LLMs for Predicting Unknown Diagnoses from Clinical Notes
by: Albassam, Dina, et al.
Published: (2025)
by: Albassam, Dina, et al.
Published: (2025)
Detecting Knowledge Boundary of Vision Large Language Models by Sampling-Based Inference
by: Chen, Zhuo, et al.
Published: (2025)
by: Chen, Zhuo, et al.
Published: (2025)
On the Landauer formula in interfacial thermal transport
by: Dai, Jinghang, et al.
Published: (2026)
by: Dai, Jinghang, et al.
Published: (2026)
Crystal-like thermal transport in amorphous carbon
by: Moon, Jaeyun, et al.
Published: (2024)
by: Moon, Jaeyun, et al.
Published: (2024)
Fire360: A Benchmark for Robust Perception and Episodic Memory in Degraded 360-Degree Firefighting Videos
by: Tiwari, Aditi, et al.
Published: (2025)
by: Tiwari, Aditi, et al.
Published: (2025)
Smart Starts: Accelerating Convergence through Uncommon Region Exploration
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Consciousness as Uncommon Self-Knowledge: A Synergistic Information Framework
by: Tallam, Krti
Published: (2026)
by: Tallam, Krti
Published: (2026)
CTLformer: A Hybrid Denoising Model Combining Convolutional Layers and Self-Attention for Enhanced CT Image Reconstruction
by: Zheng, Zhiting, et al.
Published: (2025)
by: Zheng, Zhiting, et al.
Published: (2025)
MM-Retinal V2: Transfer an Elite Knowledge Spark into Fundus Vision-Language Pretraining
by: Wu, Ruiqi, et al.
Published: (2025)
by: Wu, Ruiqi, et al.
Published: (2025)
ORBIT: Cost-Effective Dataset Curation for Large Language Model Domain Adaptation with an Astronomy Case Study
by: Modesitt, Eric, et al.
Published: (2024)
by: Modesitt, Eric, et al.
Published: (2024)
Vision-G1: Towards General Vision Language Reasoning with Multi-Domain Data Curation
by: Zha, Yuheng, et al.
Published: (2025)
by: Zha, Yuheng, et al.
Published: (2025)
Image Recognition with Vision and Language Embeddings of VLMs
by: Volkov, Illia, et al.
Published: (2025)
by: Volkov, Illia, et al.
Published: (2025)
Similar Items
-
Cache-of-Thought: Master-Apprentice Framework for Cost-Effective Vision Language Model Reasoning
by: Wu, Mingyuan, et al.
Published: (2025) -
Spatio-Temporal LLM: Reasoning about Environments and Actions
by: Zheng, Haozhen, et al.
Published: (2025) -
AquaVLM: Improving Underwater Situation Awareness with Mobile Vision Language Models
by: Tian, Beitong, et al.
Published: (2025) -
AquaScope: Reliable Underwater Image Transmission on Mobile Devices
by: Tian, Beitong, et al.
Published: (2025) -
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
by: Wu, Mingyuan, et al.
Published: (2025)