K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
Fuente:
arXiv
Salvato in:
| Autori principali: | Kim, Soyeon, Kang, Cheongwoong, Lee, Myeongjin, Chang, Eun-Chul, Lee, Jaedeok, Choi, Jaesik |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Spectral Integrated Gradients for Coarse-to-Fine Feature Attribution
di: Kim, Soyeon, et al.
Pubblicazione: (2026)
di: Kim, Soyeon, et al.
Pubblicazione: (2026)
Manifold-Aligned Guided Integrated Gradients for Reliable Feature Attribution
di: Kim, Soyeon, et al.
Pubblicazione: (2026)
di: Kim, Soyeon, et al.
Pubblicazione: (2026)
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
di: Raoufi, Behnam, et al.
Pubblicazione: (2025)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
di: Dai, Song, et al.
Pubblicazione: (2025)
di: Dai, Song, et al.
Pubblicazione: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
di: Tong, Jingqi, et al.
Pubblicazione: (2025)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
di: Cui, Shaoyang, et al.
Pubblicazione: (2026)
di: Cui, Shaoyang, et al.
Pubblicazione: (2026)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
di: Shah, Nisarg A., et al.
Pubblicazione: (2025)
di: Shah, Nisarg A., et al.
Pubblicazione: (2025)
NOAH: Benchmarking Narrative Prior driven Hallucination and Omission in Video Large Language Models
di: Lee, Kyuho, et al.
Pubblicazione: (2025)
di: Lee, Kyuho, et al.
Pubblicazione: (2025)
Universal Adversarial Attack on Aligned Multimodal LLMs
di: Rahmatullaev, Temurbek, et al.
Pubblicazione: (2025)
di: Rahmatullaev, Temurbek, et al.
Pubblicazione: (2025)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
di: Yang, Shan
Pubblicazione: (2026)
di: Yang, Shan
Pubblicazione: (2026)
GeoHeight-Bench: Towards Height-Aware Multimodal Reasoning in Remote Sensing
di: Hu, Xuran, et al.
Pubblicazione: (2026)
di: Hu, Xuran, et al.
Pubblicazione: (2026)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
di: Zhang, Sinin, et al.
Pubblicazione: (2026)
di: Zhang, Sinin, et al.
Pubblicazione: (2026)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
di: Le, Van-Truong
Pubblicazione: (2026)
di: Le, Van-Truong
Pubblicazione: (2026)
U-Net-Like Spiking Neural Networks for Single Image Dehazing
di: Li, Huibin, et al.
Pubblicazione: (2025)
di: Li, Huibin, et al.
Pubblicazione: (2025)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
di: Liu, Bingnan, et al.
Pubblicazione: (2026)
di: Liu, Bingnan, et al.
Pubblicazione: (2026)
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
di: Asanuma, Haruka, et al.
Pubblicazione: (2025)
di: Asanuma, Haruka, et al.
Pubblicazione: (2025)
GTPBD: A Fine-Grained Global Terraced Parcel and Boundary Dataset
di: Zhang, Zhiwei, et al.
Pubblicazione: (2025)
di: Zhang, Zhiwei, et al.
Pubblicazione: (2025)
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
di: Li, Danyang, et al.
Pubblicazione: (2025)
di: Li, Danyang, et al.
Pubblicazione: (2025)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
di: Romero, Angel, et al.
Pubblicazione: (2025)
di: Romero, Angel, et al.
Pubblicazione: (2025)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
di: Gasparino, Mateus Valverde, et al.
Pubblicazione: (2024)
di: Gasparino, Mateus Valverde, et al.
Pubblicazione: (2024)
EmbodiedLGR: Integrating Lightweight Graph Representation and Retrieval for Semantic-Spatial Memory in Robotic Agents
di: Riva, Paolo, et al.
Pubblicazione: (2026)
di: Riva, Paolo, et al.
Pubblicazione: (2026)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
di: Lim, Shoon Kit, et al.
Pubblicazione: (2025)
di: Lim, Shoon Kit, et al.
Pubblicazione: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
Deep Probabilistic Traversability with Test-time Adaptation for Uncertainty-aware Planetary Rover Navigation
di: Endo, Masafumi, et al.
Pubblicazione: (2024)
di: Endo, Masafumi, et al.
Pubblicazione: (2024)
CoMoCAVs: Cohesive Decision-Guided Motion Planning for Connected and Autonomous Vehicles with Multi-Policy Reinforcement Learning
di: Hu, Pan
Pubblicazione: (2025)
di: Hu, Pan
Pubblicazione: (2025)
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos
di: Li, Yayuan, et al.
Pubblicazione: (2025)
di: Li, Yayuan, et al.
Pubblicazione: (2025)
PathFormer: A Transformer with 3D Grid Constraints for Digital Twin Robot-Arm Trajectory Generation
di: Alanazi, Ahmed, et al.
Pubblicazione: (2025)
di: Alanazi, Ahmed, et al.
Pubblicazione: (2025)
Motion Perceiver: Real-Time Occupancy Forecasting for Embedded Systems
di: Ferenczi, Bryce, et al.
Pubblicazione: (2023)
di: Ferenczi, Bryce, et al.
Pubblicazione: (2023)
MemeCraft: Contextual and Stance-Driven Multimodal Meme Generation
di: Wang, Han, et al.
Pubblicazione: (2024)
di: Wang, Han, et al.
Pubblicazione: (2024)
Evaluating Visual Mathematics in Multimodal LLMs: A Multilingual Benchmark Based on the Kangaroo Tests
di: Sáez, Arnau Igualde, et al.
Pubblicazione: (2025)
di: Sáez, Arnau Igualde, et al.
Pubblicazione: (2025)
Beyond Routing: Characterising Expert Tuning and Representation in Vision Mixture-of-Experts
di: Tangtartharakul, Gene, et al.
Pubblicazione: (2026)
di: Tangtartharakul, Gene, et al.
Pubblicazione: (2026)
SUN Team's Contribution to ABAW 2024 Competition: Audio-visual Valence-Arousal Estimation and Expression Recognition
di: Dresvyanskiy, Denis, et al.
Pubblicazione: (2024)
di: Dresvyanskiy, Denis, et al.
Pubblicazione: (2024)
Learning the meanings of function words from grounded language using a visual question answering model
di: Portelance, Eva, et al.
Pubblicazione: (2023)
di: Portelance, Eva, et al.
Pubblicazione: (2023)
Closed-Loop Neural Activation Control in Vision-Language-Action Models
di: Babu, Abhijith, et al.
Pubblicazione: (2026)
di: Babu, Abhijith, et al.
Pubblicazione: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
di: Gopinathan, Muraleekrishna, et al.
Pubblicazione: (2024)
A Segmented Robot Grasping Perception Neural Network for Edge AI
di: Bröcheler, Casper, et al.
Pubblicazione: (2025)
di: Bröcheler, Casper, et al.
Pubblicazione: (2025)
VisChainBench: A Benchmark for Multi-Turn, Multi-Image Visual Reasoning Beyond Language Priors
di: Lyu, Wenbo, et al.
Pubblicazione: (2025)
di: Lyu, Wenbo, et al.
Pubblicazione: (2025)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
di: Durrani, Hamza Ahmed, et al.
Pubblicazione: (2026)
di: Durrani, Hamza Ahmed, et al.
Pubblicazione: (2026)
Vision-based Situational Graphs Exploiting Fiducial Markers for the Integration of Semantic Entities
di: Tourani, Ali, et al.
Pubblicazione: (2023)
di: Tourani, Ali, et al.
Pubblicazione: (2023)
UAV-assisted Visual SLAM Generating Reconstructed 3D Scene Graphs in GPS-denied Environments
di: Radwan, Ahmed, et al.
Pubblicazione: (2024)
di: Radwan, Ahmed, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Spectral Integrated Gradients for Coarse-to-Fine Feature Attribution
di: Kim, Soyeon, et al.
Pubblicazione: (2026) -
Manifold-Aligned Guided Integrated Gradients for Reliable Feature Attribution
di: Kim, Soyeon, et al.
Pubblicazione: (2026) -
CLIP-Joint-Detect: End-to-End Joint Training of Object Detectors with Contrastive Vision-Language Supervision
di: Raoufi, Behnam, et al.
Pubblicazione: (2025) -
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
di: Dai, Song, et al.
Pubblicazione: (2025) -
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
di: Tong, Jingqi, et al.
Pubblicazione: (2025)