Diagram-Driven Course Questions Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Zhang, Xinyu, Zhang, Lingling, Wu, Yanrui, Huang, Muye, Wu, Wenjun, Li, Bo, Wang, Shaowei, Fernando, Basura, Liu, Jun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
VProChart: Answering Chart Question through Visual Perception Alignment Agent and Programmatic Solution Reasoning
by: Huang, Muye, et al.
Published: (2024)
by: Huang, Muye, et al.
Published: (2024)
GoT-CQA: Graph-of-Thought Guided Compositional Reasoning for Chart Question Answering
by: Zhang, Lingling, et al.
Published: (2024)
by: Zhang, Lingling, et al.
Published: (2024)
EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart Understanding
by: Huang, Muye, et al.
Published: (2024)
by: Huang, Muye, et al.
Published: (2024)
SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and More
by: Huang, Muye, et al.
Published: (2026)
by: Huang, Muye, et al.
Published: (2026)
ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
by: Huang, Muye, et al.
Published: (2025)
by: Huang, Muye, et al.
Published: (2025)
Mitigating Easy Option Bias in Multiple-Choice Question Answering
by: Zhang, Hao, et al.
Published: (2025)
by: Zhang, Hao, et al.
Published: (2025)
ChartAct: A Benchmark for Dynamic Chart Understanding
by: Huang, Muye, et al.
Published: (2026)
by: Huang, Muye, et al.
Published: (2026)
CoFFT: Chain of Foresight-Focus Thought for Visual Language Models
by: Zhang, Xinyu, et al.
Published: (2025)
by: Zhang, Xinyu, et al.
Published: (2025)
Effectively Leveraging CLIP for Generating Situational Summaries of Images and Videos
by: Verma, Dhruv, et al.
Published: (2024)
by: Verma, Dhruv, et al.
Published: (2024)
Learning to Generate Long-term Future Narrations Describing Activities of Daily Living
by: Rajendiran, Ramanathan, et al.
Published: (2025)
by: Rajendiran, Ramanathan, et al.
Published: (2025)
RCA: Region Conditioned Adaptation for Visual Abductive Reasoning
by: Zhang, Hao, et al.
Published: (2023)
by: Zhang, Hao, et al.
Published: (2023)
VOST-SGG: VLM-Aided One-Stage Spatio-Temporal Scene Graph Generation
by: Sugandhika, Chinthani, et al.
Published: (2025)
by: Sugandhika, Chinthani, et al.
Published: (2025)
HMR3D: Hierarchical Multimodal Representation for 3D Scene Understanding with Large Vision-Language Model
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
Neuro Symbolic Knowledge Reasoning for Procedural Video Question Answering
by: Fernando, Basura, et al.
Published: (2025)
by: Fernando, Basura, et al.
Published: (2025)
Predicting the Next Action by Modeling the Abstract Goal
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
Interaction Region Visual Transformer for Egocentric Action Anticipation
by: Roy, Debaditya, et al.
Published: (2022)
by: Roy, Debaditya, et al.
Published: (2022)
Modelling Spatio-Temporal Interactions For Compositional Action Recognition
by: Rajendiran, Ramanathan, et al.
Published: (2023)
by: Rajendiran, Ramanathan, et al.
Published: (2023)
Neurodynamics-Driven Coupled Neural P Systems for Multi-Focus Image Fusion
by: Li, Bo, et al.
Published: (2025)
by: Li, Bo, et al.
Published: (2025)
Training-Free Action Recognition and Goal Inference with Dynamic Frame Selection
by: Keat, Ee Yeo, et al.
Published: (2024)
by: Keat, Ee Yeo, et al.
Published: (2024)
Situational Scene Graph for Structured Human-centric Situation Understanding
by: Sugandhika, Chinthani, et al.
Published: (2024)
by: Sugandhika, Chinthani, et al.
Published: (2024)
Know-Show: Benchmarking Video-Language Models on Spatio-Temporal Grounded Reasoning
by: Sugandhika, Chinthani, et al.
Published: (2025)
by: Sugandhika, Chinthani, et al.
Published: (2025)
Improving Temporal Action Segmentation via Constraint-Aware Decoding
by: Ee, Yeo Keat, et al.
Published: (2026)
by: Ee, Yeo Keat, et al.
Published: (2026)
Cross-Modal Few-Shot Learning: a Generative Transfer Learning Framework
by: Yang, Zhengwei, et al.
Published: (2024)
by: Yang, Zhengwei, et al.
Published: (2024)
ChainReaction: Causal Chain-Guided Reasoning for Modular and Explainable Causal-Why Video Question Answering
by: Parmar, Paritosh, et al.
Published: (2025)
by: Parmar, Paritosh, et al.
Published: (2025)
ReDiffuse: Rotation Equivariant Diffusion Model for Multi-focus Image Fusion
by: Li, Bo, et al.
Published: (2026)
by: Li, Bo, et al.
Published: (2026)
MagicGeo: Training-Free Text-Guided Geometric Diagram Generation
by: Wang, Junxiao, et al.
Published: (2025)
by: Wang, Junxiao, et al.
Published: (2025)
Task-Driven Prompt Learning: A Joint Framework for Multi-modal Cloud Removal and Segmentation
by: Zhang, Zaiyan, et al.
Published: (2026)
by: Zhang, Zaiyan, et al.
Published: (2026)
VoQA: Visual-only Question Answering
by: An, Jianing, et al.
Published: (2025)
by: An, Jianing, et al.
Published: (2025)
GeoLoom: High-quality Geometric Diagram Generation from Textual Input
by: Wei, Xiaojing, et al.
Published: (2025)
by: Wei, Xiaojing, et al.
Published: (2025)
Geoparsing: Diagram Parsing for Plane and Solid Geometry with a Unified Formal Language
by: Wang, Peijie, et al.
Published: (2026)
by: Wang, Peijie, et al.
Published: (2026)
Inferring Past Human Actions in Homes with Abductive Reasoning
by: Tan, Clement, et al.
Published: (2022)
by: Tan, Clement, et al.
Published: (2022)
SparrowVQE: Visual Question Explanation for Course Content Understanding
by: Li, Jialu, et al.
Published: (2024)
by: Li, Jialu, et al.
Published: (2024)
LSVOS Challenge 3rd Place Report: SAM2 and Cutie based VOS
by: Liu, Xinyu, et al.
Published: (2024)
by: Liu, Xinyu, et al.
Published: (2024)
TinyLLaVA-Video-R1: Towards Smaller LMMs for Video Reasoning
by: Zhang, Xingjian, et al.
Published: (2025)
by: Zhang, Xingjian, et al.
Published: (2025)
EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use
by: Wen, Siwei, et al.
Published: (2026)
by: Wen, Siwei, et al.
Published: (2026)
T-SVG: Text-Driven Stereoscopic Video Generation
by: Jin, Qiao, et al.
Published: (2024)
by: Jin, Qiao, et al.
Published: (2024)
SciFlow-Bench: Evaluating Structure-Aware Scientific Diagram Generation via Inverse Parsing
by: Zhang, Tong, et al.
Published: (2026)
by: Zhang, Tong, et al.
Published: (2026)
IMoRe: Implicit Program-Guided Reasoning for Human Motion Q&A
by: Li, Chen, et al.
Published: (2025)
by: Li, Chen, et al.
Published: (2025)
GenView++: Unifying Adaptive Generative Augmentation and Quality-Driven Supervision for Contrastive Representation Learning
by: Li, Xiaojie, et al.
Published: (2025)
by: Li, Xiaojie, et al.
Published: (2025)
RxnCaption: Reformulating Reaction Diagram Parsing as Visual Prompt Guided Captioning
by: Song, Jiahe, et al.
Published: (2025)
by: Song, Jiahe, et al.
Published: (2025)
Similar Items
-
VProChart: Answering Chart Question through Visual Perception Alignment Agent and Programmatic Solution Reasoning
by: Huang, Muye, et al.
Published: (2024) -
GoT-CQA: Graph-of-Thought Guided Compositional Reasoning for Chart Question Answering
by: Zhang, Lingling, et al.
Published: (2024) -
EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart Understanding
by: Huang, Muye, et al.
Published: (2024) -
SketchVL: Policy Optimization via Fine-Grained Credit Assignment for Chart Understanding and More
by: Huang, Muye, et al.
Published: (2026) -
ChartSketcher: Reasoning with Multimodal Feedback and Reflection for Chart Understanding
by: Huang, Muye, et al.
Published: (2025)