EduFlow: Advancing MLLMs' Problem-Solving Proficiency through Multi-Stage, Multi-Perspective Critique
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhu, Chenglin, Zhang, Tao, Li, Chong, Lin, Mingan, Zhou, Zenan, Xie, Jian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
von: Chen, Yuhao, et al.
Veröffentlicht: (2026)
von: Chen, Yuhao, et al.
Veröffentlicht: (2026)
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
von: Romero, Angel, et al.
Veröffentlicht: (2025)
von: Romero, Angel, et al.
Veröffentlicht: (2025)
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
von: Gasparino, Mateus Valverde, et al.
Veröffentlicht: (2024)
von: Gasparino, Mateus Valverde, et al.
Veröffentlicht: (2024)
AGOP as Explanation: From Feature Learning to Per-Sample Attribution in Image Classifiers
von: Katakam, Raj Kiran Gupta
Veröffentlicht: (2026)
von: Katakam, Raj Kiran Gupta
Veröffentlicht: (2026)
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)
SUN Team's Contribution to ABAW 2024 Competition: Audio-visual Valence-Arousal Estimation and Expression Recognition
von: Dresvyanskiy, Denis, et al.
Veröffentlicht: (2024)
von: Dresvyanskiy, Denis, et al.
Veröffentlicht: (2024)
Learning the meanings of function words from grounded language using a visual question answering model
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
von: Portelance, Eva, et al.
Veröffentlicht: (2023)
Closed-Loop Neural Activation Control in Vision-Language-Action Models
von: Babu, Abhijith, et al.
Veröffentlicht: (2026)
von: Babu, Abhijith, et al.
Veröffentlicht: (2026)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
von: Yang, Shan
Veröffentlicht: (2026)
von: Yang, Shan
Veröffentlicht: (2026)
QoSGMAA: A Robust Multi-Order Graph Attention and Adversarial Framework for Sparse QoS Prediction
von: Du, Guanchen, et al.
Veröffentlicht: (2025)
von: Du, Guanchen, et al.
Veröffentlicht: (2025)
SpatialMath: Spatial Comprehension-Infused Symbolic Reasoning for Mathematical Problem-Solving
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026)
von: Bajpai, Ashutosh, et al.
Veröffentlicht: (2026)
CLARE: Continual Learning for Vision-Language-Action Models via Autonomous Adapter Routing and Expansion
von: Römer, Ralf, et al.
Veröffentlicht: (2026)
von: Römer, Ralf, et al.
Veröffentlicht: (2026)
Cortex 2.0: Grounding World Models in Real-World Industrial Deployment
von: Aida, Adriana, et al.
Veröffentlicht: (2026)
von: Aida, Adriana, et al.
Veröffentlicht: (2026)
ExpReS-VLA: Specializing Vision-Language-Action Models Through Experience Replay and Retrieval
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
von: Syed, Shahram Najam, et al.
Veröffentlicht: (2025)
eStonefish-Scenes: A Sim-to-Real Validated and Robot-Centric Event-based Optical Flow Dataset for Underwater Vehicles
von: Mansour, Jad, et al.
Veröffentlicht: (2025)
von: Mansour, Jad, et al.
Veröffentlicht: (2025)
VLM-VPI: A Vision-Language Reasoning Framework for Improving Automated Vehicle-Pedestrian Interactions
von: Pu, Qingwen, et al.
Veröffentlicht: (2026)
von: Pu, Qingwen, et al.
Veröffentlicht: (2026)
Survey Transfer Learning: Recycling Data with Silicon Responses
von: Amini, Ali
Veröffentlicht: (2025)
von: Amini, Ali
Veröffentlicht: (2025)
Mitigating Catastrophic Forgetting in the Incremental Learning of Medical Images
von: Yavari, Sara, et al.
Veröffentlicht: (2025)
von: Yavari, Sara, et al.
Veröffentlicht: (2025)
ReLKD: Inter-Class Relation Learning with Knowledge Distillation for Generalized Category Discovery
von: Zhou, Fang, et al.
Veröffentlicht: (2025)
von: Zhou, Fang, et al.
Veröffentlicht: (2025)
Efficient Image Pre-Training with Siamese Cropped Masked Autoencoders
von: Eymaël, Alexandre, et al.
Veröffentlicht: (2024)
von: Eymaël, Alexandre, et al.
Veröffentlicht: (2024)
Reducing the Sensitivity of Neural Physics Simulators to Mesh Topology via Pretraining
von: Vaska, Nathan, et al.
Veröffentlicht: (2025)
von: Vaska, Nathan, et al.
Veröffentlicht: (2025)
Data Organization Matters in Multimodal Instruction Tuning: A Controlled Study of Capability Trade-offs
von: Tang, Guowei
Veröffentlicht: (2026)
von: Tang, Guowei
Veröffentlicht: (2026)
RDPO: Real Data Preference Optimization for Physics Consistency Video Generation
von: Qian, Wenxu, et al.
Veröffentlicht: (2025)
von: Qian, Wenxu, et al.
Veröffentlicht: (2025)
CrystalDiT: A Diffusion Transformer for Crystal Generation
von: Yi, Xiaohan, et al.
Veröffentlicht: (2025)
von: Yi, Xiaohan, et al.
Veröffentlicht: (2025)
Adaptive Self-Training for Object Detection
von: Vandeghen, Renaud, et al.
Veröffentlicht: (2022)
von: Vandeghen, Renaud, et al.
Veröffentlicht: (2022)
Once-For-All: A Train-Once and Select-Anytime Framework for Multimodal Instruction Tuning
von: Dong, Mingkang, et al.
Veröffentlicht: (2026)
von: Dong, Mingkang, et al.
Veröffentlicht: (2026)
From Demonstrations to Safe Deployment: Path-Consistent Safety Filtering for Diffusion Policies
von: Römer, Ralf, et al.
Veröffentlicht: (2025)
von: Römer, Ralf, et al.
Veröffentlicht: (2025)
Curb Your Attention: Causal Attention Gating for Robust Trajectory Prediction in Autonomous Driving
von: Ahmadi, Ehsan, et al.
Veröffentlicht: (2024)
von: Ahmadi, Ehsan, et al.
Veröffentlicht: (2024)
OnlyFlow: Optical Flow based Motion Conditioning for Video Diffusion Models
von: Koroglu, Mathis, et al.
Veröffentlicht: (2024)
von: Koroglu, Mathis, et al.
Veröffentlicht: (2024)
A Comparative Survey of PyTorch vs TensorFlow for Deep Learning: Usability, Performance, and Deployment Trade-offs
von: Alawi, Zakariya Ba
Veröffentlicht: (2025)
von: Alawi, Zakariya Ba
Veröffentlicht: (2025)
FedWCM: Unleashing the Potential of Momentum-based Federated Learning in Long-Tailed Scenarios
von: Li, Tianle, et al.
Veröffentlicht: (2025)
von: Li, Tianle, et al.
Veröffentlicht: (2025)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
von: Portelance, Eva, et al.
Veröffentlicht: (2024)
von: Portelance, Eva, et al.
Veröffentlicht: (2024)
Perceptual Flow Network for Visually Grounded Reasoning
von: Li, Yangfu, et al.
Veröffentlicht: (2026)
von: Li, Yangfu, et al.
Veröffentlicht: (2026)
Invariant Representation via Decoupling Style and Spurious Features from Images
von: Li, Ruimeng, et al.
Veröffentlicht: (2023)
von: Li, Ruimeng, et al.
Veröffentlicht: (2023)
eCARLA-scenes: A synthetically generated dataset for event-based optical flow prediction
von: Mansour, Jad, et al.
Veröffentlicht: (2024)
von: Mansour, Jad, et al.
Veröffentlicht: (2024)
CourseTimeQA: A Lecture-Video Benchmark and a Latency-Constrained Cross-Modal Fusion Method for Timestamped QA
von: Kovalev, Vsevolod, et al.
Veröffentlicht: (2025)
von: Kovalev, Vsevolod, et al.
Veröffentlicht: (2025)
Selection, Not Fusion: Radar-Modulated State Space Models for Radar-Camera Depth Estimation
von: Hou, Zhangcheng, et al.
Veröffentlicht: (2026)
von: Hou, Zhangcheng, et al.
Veröffentlicht: (2026)
Exposing Blindspots: Cultural Bias Evaluation in Generative Image Models
von: Seo, Huichan, et al.
Veröffentlicht: (2025)
von: Seo, Huichan, et al.
Veröffentlicht: (2025)
Pointing-Based Object Recognition
von: Hajdúch, Lukáš, et al.
Veröffentlicht: (2026)
von: Hajdúch, Lukáš, et al.
Veröffentlicht: (2026)
Smooth regularization for efficient video recognition
von: Goldman, Gil, et al.
Veröffentlicht: (2025)
von: Goldman, Gil, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
EvoPrune: Early-Stage Visual Token Pruning for Efficient MLLMs
von: Chen, Yuhao, et al.
Veröffentlicht: (2026) -
Dream to Fly: Model-Based Reinforcement Learning for Vision-Based Drone Flight
von: Romero, Angel, et al.
Veröffentlicht: (2025) -
WayFASTER: a Self-Supervised Traversability Prediction for Increased Navigation Awareness
von: Gasparino, Mateus Valverde, et al.
Veröffentlicht: (2024) -
AGOP as Explanation: From Feature Learning to Per-Sample Attribution in Image Classifiers
von: Katakam, Raj Kiran Gupta
Veröffentlicht: (2026) -
Lifelong Learning in Vision-Language Models: Enhanced EWC with Cross-Modal Knowledge Retention
von: Durrani, Hamza Ahmed, et al.
Veröffentlicht: (2026)