AnE: Pushing the Reasoning Frontier of Multimodal LLMs via Anchor Evolution
Fuente:
arXiv
Guardado en:
| Autores principales: | Wang, Zehao, Zeng, Yihan, Gong, Zidong, Guo, Yuanfan, Zhu, Feng, Zhang, Hongzhi, Zhang, Wei, Zuo, Wangmeng |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
FedSmoothLoRA: Toward Smoother and Faster Convergence in Federated Low-Rank Adaptation
por: Wang, Zehao, et al.
Publicado: (2026)
por: Wang, Zehao, et al.
Publicado: (2026)
GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration
por: Xu, Wan, et al.
Publicado: (2025)
por: Xu, Wan, et al.
Publicado: (2025)
OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization
por: Zhu, Feng, et al.
Publicado: (2026)
por: Zhu, Feng, et al.
Publicado: (2026)
Self-Supervised Learning for Real-World Super-Resolution from Dual and Multiple Zoomed Observations
por: Zhang, Zhilu, et al.
Publicado: (2024)
por: Zhang, Zhilu, et al.
Publicado: (2024)
PhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration Synthesis
por: Yang, Yu, et al.
Publicado: (2025)
por: Yang, Yu, et al.
Publicado: (2025)
MasterWeaver: Taming Editability and Face Identity for Personalized Text-to-Image Generation
por: Wei, Yuxiang, et al.
Publicado: (2024)
por: Wei, Yuxiang, et al.
Publicado: (2024)
FramePainter: Endowing Interactive Image Editing with Video Diffusion Priors
por: Zhang, Yabo, et al.
Publicado: (2025)
por: Zhang, Yabo, et al.
Publicado: (2025)
Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning
por: Ni, Minheng, et al.
Publicado: (2024)
por: Ni, Minheng, et al.
Publicado: (2024)
MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM
por: Dong, Bowen, et al.
Publicado: (2025)
por: Dong, Bowen, et al.
Publicado: (2025)
Tool-R1: Sample-Efficient Reinforcement Learning for Agentic Tool Use
por: Zhang, Yabo, et al.
Publicado: (2025)
por: Zhang, Yabo, et al.
Publicado: (2025)
UniRestorer: Universal Image Restoration via Adaptively Estimating Image Degradation at Proper Granularity
por: Lin, Jingbo, et al.
Publicado: (2024)
por: Lin, Jingbo, et al.
Publicado: (2024)
Pseudo-Label Guided Real-World Image De-weathering: A Learning Framework with Imperfect Supervision
por: Xu, Heming, et al.
Publicado: (2025)
por: Xu, Heming, et al.
Publicado: (2025)
VitaGlyph: Vitalizing Artistic Typography with Flexible Dual-branch Diffusion Models
por: Feng, Kailai, et al.
Publicado: (2024)
por: Feng, Kailai, et al.
Publicado: (2024)
Reflection Anchors for Propagation-Aware Visual Retention in Long-Chain Multimodal Reasoning
por: Gong, Xuan, et al.
Publicado: (2026)
por: Gong, Xuan, et al.
Publicado: (2026)
CGL: Advancing Continual GUI Learning via Reinforcement Fine-Tuning
por: Yao, Zhenquan, et al.
Publicado: (2026)
por: Yao, Zhenquan, et al.
Publicado: (2026)
DreamPhysics: Learning Physics-Based 3D Dynamics with Video Diffusion Priors
por: Huang, Tianyu, et al.
Publicado: (2024)
por: Huang, Tianyu, et al.
Publicado: (2024)
PhysChoreo: Physics-Controllable Video Generation with Part-Aware Semantic Grounding
por: Zhang, Haoze, et al.
Publicado: (2025)
por: Zhang, Haoze, et al.
Publicado: (2025)
Premier: Personalized Preference Modulation with Learnable User Embedding in Text-to-Image Generation
por: Wang, Zihao, et al.
Publicado: (2026)
por: Wang, Zihao, et al.
Publicado: (2026)
SpatialGeo:Boosting Spatial Reasoning in Multimodal LLMs via Geometry-Semantics Fusion
por: Guo, Jiajie, et al.
Publicado: (2025)
por: Guo, Jiajie, et al.
Publicado: (2025)
Pushing the Frontier of Audiovisual Perception with Large-Scale Multimodal Correspondence Learning
por: Vyas, Apoorv, et al.
Publicado: (2025)
por: Vyas, Apoorv, et al.
Publicado: (2025)
ConSept: Continual Semantic Segmentation via Adapter-based Vision Transformer
por: Dong, Bowen, et al.
Publicado: (2024)
por: Dong, Bowen, et al.
Publicado: (2024)
Thin-Plate Spline-based Interpolation for Animation Line Inbetweening
por: Zhu, Tianyi, et al.
Publicado: (2024)
por: Zhu, Tianyi, et al.
Publicado: (2024)
AdaTSQ: Pushing the Pareto Frontier of Diffusion Transformers via Temporal-Sensitivity Quantization
por: Zhang, Shaoqiu, et al.
Publicado: (2026)
por: Zhang, Shaoqiu, et al.
Publicado: (2026)
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning
por: Ni, Minheng, et al.
Publicado: (2025)
por: Ni, Minheng, et al.
Publicado: (2025)
CoEditor++: Instruction-based Visual Editing via Cognitive Reasoning
por: Ni, Minheng, et al.
Publicado: (2026)
por: Ni, Minheng, et al.
Publicado: (2026)
Rethinking Transformer-Based Blind-Spot Network for Self-Supervised Image Denoising
por: Li, Junyi, et al.
Publicado: (2024)
por: Li, Junyi, et al.
Publicado: (2024)
Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers
por: Su, Zhaochen, et al.
Publicado: (2025)
por: Su, Zhaochen, et al.
Publicado: (2025)
Aggregating Nearest Sharp Features via Hybrid Transformers for Video Deblurring
por: Shang, Wei, et al.
Publicado: (2023)
por: Shang, Wei, et al.
Publicado: (2023)
DreamControl: Control-Based Text-to-3D Generation with 3D Self-Prior
por: Huang, Tianyu, et al.
Publicado: (2023)
por: Huang, Tianyu, et al.
Publicado: (2023)
EMOv2: Pushing 5M Vision Model Frontier
por: Zhang, Jiangning, et al.
Publicado: (2024)
por: Zhang, Jiangning, et al.
Publicado: (2024)
Anchor3DLane++: 3D Lane Detection via Sample-Adaptive Sparse 3D Anchor Regression
por: Huang, Shaofei, et al.
Publicado: (2024)
por: Huang, Shaofei, et al.
Publicado: (2024)
TextField3D: Towards Enhancing Open-Vocabulary 3D Generation with Noisy Text Fields
por: Huang, Tianyu, et al.
Publicado: (2023)
por: Huang, Tianyu, et al.
Publicado: (2023)
Bridging Geometry-Coherent Text-to-3D Generation with Multi-View Diffusion Priors and Gaussian Splatting
por: Yang, Feng, et al.
Publicado: (2025)
por: Yang, Feng, et al.
Publicado: (2025)
Responsible Visual Editing
por: Ni, Minheng, et al.
Publicado: (2024)
por: Ni, Minheng, et al.
Publicado: (2024)
Image Demoiréing Using Dual Camera Fusion on Mobile Phones
por: Mei, Yanting, et al.
Publicado: (2025)
por: Mei, Yanting, et al.
Publicado: (2025)
NIR-Assisted Image Denoising: A Selective Fusion Approach and A Real-World Benchmark Dataset
por: Xu, Rongjian, et al.
Publicado: (2024)
por: Xu, Rongjian, et al.
Publicado: (2024)
Dual-Camera Smooth Zoom on Mobile Phones
por: Wu, Renlong, et al.
Publicado: (2024)
por: Wu, Renlong, et al.
Publicado: (2024)
VTimeCoT: Thinking by Drawing for Video Temporal Grounding and Reasoning
por: Zhang, Jinglei, et al.
Publicado: (2025)
por: Zhang, Jinglei, et al.
Publicado: (2025)
DC-Reg: Globally Optimal Point Cloud Registration via Tight Bounding with Difference of Convex Programming
por: Lian, Wei, et al.
Publicado: (2026)
por: Lian, Wei, et al.
Publicado: (2026)
PanGu-Draw: Advancing Resource-Efficient Text-to-Image Synthesis with Time-Decoupled Training and Reusable Coop-Diffusion
por: Lu, Guansong, et al.
Publicado: (2023)
por: Lu, Guansong, et al.
Publicado: (2023)
Ejemplares similares
-
FedSmoothLoRA: Toward Smoother and Faster Convergence in Federated Low-Rank Adaptation
por: Wang, Zehao, et al.
Publicado: (2026) -
GLaVE-Cap: Global-Local Aligned Video Captioning with Vision Expert Integration
por: Xu, Wan, et al.
Publicado: (2025) -
OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization
por: Zhu, Feng, et al.
Publicado: (2026) -
Self-Supervised Learning for Real-World Super-Resolution from Dual and Multiple Zoomed Observations
por: Zhang, Zhilu, et al.
Publicado: (2024) -
PhysWorld: From Real Videos to World Models of Deformable Objects via Physics-Aware Demonstration Synthesis
por: Yang, Yu, et al.
Publicado: (2025)