FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding
Fuente:
arXiv
Saved in:
| Main Authors: | Faure, Gueter Josmy, Chen, Min-Hung, Yeh, Jia-Fong, Su, Hung-Ting, Hsu, Winston H. |
|---|---|
| Format: | Preprint |
| Published: |
2026
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics
by: Faure, Gueter Josmy, et al.
Published: (2024)
by: Faure, Gueter Josmy, et al.
Published: (2024)
SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning
by: Chunhachatrachai, Pawat, et al.
Published: (2026)
by: Chunhachatrachai, Pawat, et al.
Published: (2026)
SceneFunRI: Reasoning the Invisible for Task-Driven Functional Object Localization
by: Chen, Posheng, et al.
Published: (2026)
by: Chen, Posheng, et al.
Published: (2026)
MovieCORE: COgnitive REasoning in Movies
by: Faure, Gueter Josmy, et al.
Published: (2025)
by: Faure, Gueter Josmy, et al.
Published: (2025)
VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions
by: Su, Hung-Ting, et al.
Published: (2026)
by: Su, Hung-Ting, et al.
Published: (2026)
ADAPT: Benchmarking Commonsense Planning under Unspecified Affordance Constraints
by: Chen, Pei-An, et al.
Published: (2026)
by: Chen, Pei-An, et al.
Published: (2026)
MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
by: Hong, Wenyi, et al.
Published: (2025)
by: Hong, Wenyi, et al.
Published: (2025)
Affordance-Guided Coarse-to-Fine Exploration for Base Placement in Open-Vocabulary Mobile Manipulation
by: Lin, Tzu-Jung, et al.
Published: (2025)
by: Lin, Tzu-Jung, et al.
Published: (2025)
TemporalBench: Benchmarking Fine-grained Temporal Understanding for Multimodal Video Models
by: Cai, Mu, et al.
Published: (2024)
by: Cai, Mu, et al.
Published: (2024)
Investigating Video Reasoning Capability of Large Language Models with Tropes in Movies
by: Su, Hung-Ting, et al.
Published: (2024)
by: Su, Hung-Ting, et al.
Published: (2024)
Tracking-Assisted Object Detection with Event Cameras
by: Yen, Ting-Kang, et al.
Published: (2024)
by: Yen, Ting-Kang, et al.
Published: (2024)
UniFine: A Unified and Fine-grained Approach for Zero-shot Vision-Language Understanding
by: Wang, Zhecan, et al.
Published: (2023)
by: Wang, Zhecan, et al.
Published: (2023)
CT-FineBench: A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report Generation
by: Yuan, Ruifeng, et al.
Published: (2026)
by: Yuan, Ruifeng, et al.
Published: (2026)
MA-Bench: Towards Fine-grained Micro-Action Understanding
by: Li, Kun, et al.
Published: (2026)
by: Li, Kun, et al.
Published: (2026)
ViGoEmotions: A Benchmark Dataset For Fine-grained Emotion Detection on Vietnamese Texts
by: Tran, Hung Quang, et al.
Published: (2026)
by: Tran, Hung Quang, et al.
Published: (2026)
MMDocBench: Benchmarking Large Vision-Language Models for Fine-Grained Visual Document Understanding
by: Zhu, Fengbin, et al.
Published: (2024)
by: Zhu, Fengbin, et al.
Published: (2024)
Unveiling Narrative Reasoning Limits of Large Language Models with Trope in Movie Synopses
by: Su, Hung-Ting, et al.
Published: (2024)
by: Su, Hung-Ting, et al.
Published: (2024)
The BiGGen Bench: A Principled Benchmark for Fine-grained Evaluation of Language Models with Language Models
by: Kim, Seungone, et al.
Published: (2024)
by: Kim, Seungone, et al.
Published: (2024)
FineState-Bench: Benchmarking State-Conditioned Grounding for Fine-grained GUI State Setting
by: Ji, Fengxian, et al.
Published: (2026)
by: Ji, Fengxian, et al.
Published: (2026)
KPM-Bench: A Kinematic Parsing Motion Benchmark for Fine-grained Motion-centric Video Understanding
by: Lin, Boda, et al.
Published: (2026)
by: Lin, Boda, et al.
Published: (2026)
Exploring the Capabilities of LLMs for IMU-based Fine-grained Human Activity Understanding
by: Xu, Lilin, et al.
Published: (2025)
by: Xu, Lilin, et al.
Published: (2025)
RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
by: Chen, Yelin, et al.
Published: (2026)
by: Chen, Yelin, et al.
Published: (2026)
Tel2Veh: Fusion of Telecom Data and Vehicle Flow to Predict Camera-Free Traffic via a Spatio-Temporal Framework
by: Lin, ChungYi, et al.
Published: (2024)
by: Lin, ChungYi, et al.
Published: (2024)
Revisiting Semi-supervised Adversarial Robustness via Noise-aware Online Robust Distillation
by: Wu, Tsung-Han, et al.
Published: (2024)
by: Wu, Tsung-Han, et al.
Published: (2024)
FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
by: Jiang, Yuxin, et al.
Published: (2023)
by: Jiang, Yuxin, et al.
Published: (2023)
Lost in Space: Probing Fine-grained Spatial Understanding in Vision and Language Resamplers
by: Pantazopoulos, Georgios, et al.
Published: (2024)
by: Pantazopoulos, Georgios, et al.
Published: (2024)
Distribution Discrepancy and Feature Heterogeneity for Active 3D Object Detection
by: Chen, Huang-Yu, et al.
Published: (2024)
by: Chen, Huang-Yu, et al.
Published: (2024)
VICtoR: Learning Hierarchical Vision-Instruction Correlation Rewards for Long-horizon Manipulation
by: Hung, Kuo-Han, et al.
Published: (2024)
by: Hung, Kuo-Han, et al.
Published: (2024)
Large-scale and Fine-grained Vision-language Pre-training for Enhanced CT Image Understanding
by: Shui, Zhongyi, et al.
Published: (2025)
by: Shui, Zhongyi, et al.
Published: (2025)
Context-Aware Replanning with Pre-explored Semantic Map for Object Navigation
by: Ko, Po-Chen, et al.
Published: (2024)
by: Ko, Po-Chen, et al.
Published: (2024)
PhysBench: Benchmarking and Enhancing Vision-Language Models for Physical World Understanding
by: Chow, Wei, et al.
Published: (2025)
by: Chow, Wei, et al.
Published: (2025)
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
by: Sun, Boyuan, et al.
Published: (2026)
by: Sun, Boyuan, et al.
Published: (2026)
Zoom-Zero: Reinforced Coarse-to-Fine Video Understanding via Temporal Zoom-in
by: Shen, Xiaoqian, et al.
Published: (2025)
by: Shen, Xiaoqian, et al.
Published: (2025)
Analyzing Fine-Grained Alignment and Enhancing Vision Understanding in Multimodal Language Models
by: Jiang, Jiachen, et al.
Published: (2025)
by: Jiang, Jiachen, et al.
Published: (2025)
IFShip: Interpretable Fine-grained Ship Classification with Domain Knowledge-Enhanced Vision-Language Models
by: Guo, Mingning, et al.
Published: (2024)
by: Guo, Mingning, et al.
Published: (2024)
Navigating the Nuances: A Fine-grained Evaluation of Vision-Language Navigation
by: Wang, Zehao, et al.
Published: (2024)
by: Wang, Zehao, et al.
Published: (2024)
FashionFAE: Fine-grained Attributes Enhanced Fashion Vision-Language Pre-training
by: Huang, Jiale, et al.
Published: (2024)
by: Huang, Jiale, et al.
Published: (2024)
ChartHal: A Fine-grained Framework Evaluating Hallucination of Large Vision Language Models in Chart Understanding
by: Wang, Xingqi, et al.
Published: (2025)
by: Wang, Xingqi, et al.
Published: (2025)
Distilling Fine-grained Sentiment Understanding from Large Language Models
by: Zhang, Yice, et al.
Published: (2024)
by: Zhang, Yice, et al.
Published: (2024)
FActBench: A Benchmark for Fine-grained Automatic Evaluation of LLM-Generated Text in the Medical Domain
by: Afzal, Anum, et al.
Published: (2025)
by: Afzal, Anum, et al.
Published: (2025)
Similar Items
-
HERMES: temporal-coHERent long-forM understanding with Episodes and Semantics
by: Faure, Gueter Josmy, et al.
Published: (2024) -
SPATIOROUTE: Dynamic Prompt Routing for Zero-Shot Spatial Reasoning
by: Chunhachatrachai, Pawat, et al.
Published: (2026) -
SceneFunRI: Reasoning the Invisible for Task-Driven Functional Object Localization
by: Chen, Posheng, et al.
Published: (2026) -
MovieCORE: COgnitive REasoning in Movies
by: Faure, Gueter Josmy, et al.
Published: (2025) -
VLN-NF: Feasibility-Aware Vision-and-Language Navigation with False-Premise Instructions
by: Su, Hung-Ting, et al.
Published: (2026)