Gespeichert in:
| 1. Verfasser: | Dixit, Aradhya |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2601.11637 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semantic Event Graphs for Long-Form Video Question Answering
von: Dixit, Aradhya, et al.
Veröffentlicht: (2026)
von: Dixit, Aradhya, et al.
Veröffentlicht: (2026)
Curvy: A Parametric Cross-section based Surface Reconstruction
von: Mathur, Aradhya N., et al.
Veröffentlicht: (2024)
von: Mathur, Aradhya N., et al.
Veröffentlicht: (2024)
MVGaussian: High-Fidelity text-to-3D Content Generation with Multi-View Guidance and Surface Densification
von: Pham, Phu, et al.
Veröffentlicht: (2024)
von: Pham, Phu, et al.
Veröffentlicht: (2024)
NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
von: Peng, Jierui, et al.
Veröffentlicht: (2025)
von: Peng, Jierui, et al.
Veröffentlicht: (2025)
AR as an Evaluation Playground: Bridging Metrics and Visual Perception of Computer Vision Models
von: Ganj, Ashkan, et al.
Veröffentlicht: (2025)
von: Ganj, Ashkan, et al.
Veröffentlicht: (2025)
Assessing Annotation Accuracy in Ice Sheets Using Quantitative Metrics
von: Tama, Bayu Adhi, et al.
Veröffentlicht: (2024)
von: Tama, Bayu Adhi, et al.
Veröffentlicht: (2024)
HIRE: Lightweight High-Resolution Image Feature Enrichment for Multimodal LLMs
von: SR, Nikitha, et al.
Veröffentlicht: (2025)
von: SR, Nikitha, et al.
Veröffentlicht: (2025)
CorNav: Autonomous Agent with Self-Corrected Planning for Zero-Shot Vision-and-Language Navigation
von: Liang, Xiwen, et al.
Veröffentlicht: (2023)
von: Liang, Xiwen, et al.
Veröffentlicht: (2023)
Marmot: Object-Level Self-Correction via Multi-Agent Reasoning
von: Sun, Jiayang, et al.
Veröffentlicht: (2025)
von: Sun, Jiayang, et al.
Veröffentlicht: (2025)
The Essence of Balance for Self-Improving Agents in Vision-and-Language Navigation
von: Liu, Zhen, et al.
Veröffentlicht: (2026)
von: Liu, Zhen, et al.
Veröffentlicht: (2026)
A Self-Correcting Vision-Language-Action Model for Fast and Slow System Manipulation
von: Li, Chenxuan, et al.
Veröffentlicht: (2024)
von: Li, Chenxuan, et al.
Veröffentlicht: (2024)
Sherlock: Self-Correcting Reasoning in Vision-Language Models
von: Ding, Yi, et al.
Veröffentlicht: (2025)
von: Ding, Yi, et al.
Veröffentlicht: (2025)
Quantitative Metrics for Benchmarking Medical Image Harmonization
von: Parida, Abhijeet, et al.
Veröffentlicht: (2024)
von: Parida, Abhijeet, et al.
Veröffentlicht: (2024)
An Atmospheric Correction Integrated LULC Segmentation Model for High-Resolution Satellite Imagery
von: Mukherjee, Soham, et al.
Veröffentlicht: (2024)
von: Mukherjee, Soham, et al.
Veröffentlicht: (2024)
Self-Correcting Decoding with Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
von: Zhang, Ce, et al.
Veröffentlicht: (2025)
Self-Correction Inside the Model: Leveraging Layer Attention to Mitigate Hallucinations in Large Vision Language Models
von: Fu, April
Veröffentlicht: (2026)
von: Fu, April
Veröffentlicht: (2026)
PersonaVlog: Personalized Multimodal Vlog Generation with Multi-Agent Collaboration and Iterative Self-Correction
von: Hou, Xiaolu, et al.
Veröffentlicht: (2025)
von: Hou, Xiaolu, et al.
Veröffentlicht: (2025)
DepthLM: Metric Depth From Vision Language Models
von: Cai, Zhipeng, et al.
Veröffentlicht: (2025)
von: Cai, Zhipeng, et al.
Veröffentlicht: (2025)
Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases
von: Chen, Kai, et al.
Veröffentlicht: (2024)
von: Chen, Kai, et al.
Veröffentlicht: (2024)
Exploring Metric Fusion for Evaluation of NeRFs
von: Shivakumara, Shreyas, et al.
Veröffentlicht: (2025)
von: Shivakumara, Shreyas, et al.
Veröffentlicht: (2025)
Learning Self-Correction in Vision-Language Models via Rollout Augmentation
von: Ding, Yi, et al.
Veröffentlicht: (2026)
von: Ding, Yi, et al.
Veröffentlicht: (2026)
Evaluating the Evaluators: Metrics for Compositional Text-to-Image Generation
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
von: Kasaei, Seyed Amir, et al.
Veröffentlicht: (2025)
Rényi Entropy: A New Token Pruning Metric for Vision Transformers
von: Su, Wei-Yuan, et al.
Veröffentlicht: (2026)
von: Su, Wei-Yuan, et al.
Veröffentlicht: (2026)
CorrectNav: Self-Correction Flywheel Empowers Vision-Language-Action Navigation Model
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
von: Yu, Zhuoyuan, et al.
Veröffentlicht: (2025)
FD-Vision Mamba for Endoscopic Exposure Correction
von: Zheng, Zhuoran, et al.
Veröffentlicht: (2024)
von: Zheng, Zhuoran, et al.
Veröffentlicht: (2024)
PISA-Bench: The PISA Index as a Multilingual and Multimodal Metric for the Evaluation of Vision-Language Models
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
von: Haller, Patrick, et al.
Veröffentlicht: (2025)
Illumination Histogram Consistency Metric for Quantitative Assessment of Video Sequences
von: Chen, Long, et al.
Veröffentlicht: (2024)
von: Chen, Long, et al.
Veröffentlicht: (2024)
PatchEX: High-Quality Real-Time Temporal Supersampling through Patch-based Parallel Extrapolation
von: Dixit, Akanksha, et al.
Veröffentlicht: (2024)
von: Dixit, Akanksha, et al.
Veröffentlicht: (2024)
A Meaningful Perturbation Metric for Evaluating Explainability Methods
von: Cohen, Danielle, et al.
Veröffentlicht: (2025)
von: Cohen, Danielle, et al.
Veröffentlicht: (2025)
Multimodal Lengthy Videos Retrieval Framework and Evaluation Metric
von: Eltahir, Mohamed, et al.
Veröffentlicht: (2025)
von: Eltahir, Mohamed, et al.
Veröffentlicht: (2025)
Metric for Evaluating Performance of Reference-Free Demorphing Methods
von: Shukla, Nitish, et al.
Veröffentlicht: (2025)
von: Shukla, Nitish, et al.
Veröffentlicht: (2025)
HICEScore: A Hierarchical Metric for Image Captioning Evaluation
von: Zeng, Zequn, et al.
Veröffentlicht: (2024)
von: Zeng, Zequn, et al.
Veröffentlicht: (2024)
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
von: Ashraf, Tajamul, et al.
Veröffentlicht: (2025)
Self-Supervised Weighted Image Guided Quantitative MRI Super-Resolution
von: Samadifardheris, Alireza, et al.
Veröffentlicht: (2025)
von: Samadifardheris, Alireza, et al.
Veröffentlicht: (2025)
Language as Prior, Vision as Calibration: Metric Scale Recovery for Monocular Depth Estimation
von: Zhan, Mingxia, et al.
Veröffentlicht: (2026)
von: Zhan, Mingxia, et al.
Veröffentlicht: (2026)
Configurable Fairness: Direct Optimization of Parity Metrics via Vision-Language Models
von: Zhang, Miao, et al.
Veröffentlicht: (2024)
von: Zhang, Miao, et al.
Veröffentlicht: (2024)
Agent0-VL: Exploring Self-Evolving Agent for Tool-Integrated Vision-Language Reasoning
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
von: Liu, Jiaqi, et al.
Veröffentlicht: (2025)
A Quantitative Evaluation Framework for Explainable AI in Semantic Segmentation
von: Hammoud, Reem, et al.
Veröffentlicht: (2025)
von: Hammoud, Reem, et al.
Veröffentlicht: (2025)
WorldMemArena: Evaluating Multimodal Agent Memory Through Action-World Interaction
von: Liu, Chengzhi, et al.
Veröffentlicht: (2026)
von: Liu, Chengzhi, et al.
Veröffentlicht: (2026)
Self-Paced and Self-Corrective Masked Prediction for Movie Trailer Generation
von: Zhu, Sidan, et al.
Veröffentlicht: (2025)
von: Zhu, Sidan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Semantic Event Graphs for Long-Form Video Question Answering
von: Dixit, Aradhya, et al.
Veröffentlicht: (2026) -
Curvy: A Parametric Cross-section based Surface Reconstruction
von: Mathur, Aradhya N., et al.
Veröffentlicht: (2024) -
MVGaussian: High-Fidelity text-to-3D Content Generation with Multi-View Guidance and Surface Densification
von: Pham, Phu, et al.
Veröffentlicht: (2024) -
NEBULA: Do We Evaluate Vision-Language-Action Agents Correctly?
von: Peng, Jierui, et al.
Veröffentlicht: (2025) -
AR as an Evaluation Playground: Bridging Metrics and Visual Perception of Computer Vision Models
von: Ganj, Ashkan, et al.
Veröffentlicht: (2025)