Perception in Reflection
Fuente:
arXiv
Salvato in:
| Autori principali: | Wei, Yana, Zhao, Liang, Lin, Kangheng, Yu, En, Peng, Yuang, Dong, Runpei, Sun, Jianjian, Wei, Haoran, Ge, Zheng, Zhang, Xiangyu, Patel, Vishal M. |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Perception-R1: Pioneering Perception Policy with Reinforcement Learning
di: Yu, En, et al.
Pubblicazione: (2025)
di: Yu, En, et al.
Pubblicazione: (2025)
Unhackable Temporal Rewarding for Scalable Video MLLMs
di: Yu, En, et al.
Pubblicazione: (2025)
di: Yu, En, et al.
Pubblicazione: (2025)
Exploring Recurrent Long-term Temporal Fusion for Multi-view 3D Perception
di: Han, Chunrui, et al.
Pubblicazione: (2023)
di: Han, Chunrui, et al.
Pubblicazione: (2023)
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
di: Wei, Yana, et al.
Pubblicazione: (2025)
di: Wei, Yana, et al.
Pubblicazione: (2025)
Small Language Model Meets with Reinforced Vision Vocabulary
di: Wei, Haoran, et al.
Pubblicazione: (2024)
di: Wei, Haoran, et al.
Pubblicazione: (2024)
Slow Perception: Let's Perceive Geometric Figures Step-by-step
di: Wei, Haoran, et al.
Pubblicazione: (2024)
di: Wei, Haoran, et al.
Pubblicazione: (2024)
DreamLLM: Synergistic Multimodal Comprehension and Creation
di: Dong, Runpei, et al.
Pubblicazione: (2023)
di: Dong, Runpei, et al.
Pubblicazione: (2023)
General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
di: Wei, Haoran, et al.
Pubblicazione: (2024)
di: Wei, Haoran, et al.
Pubblicazione: (2024)
Merlin:Empowering Multimodal LLMs with Foresight Minds
di: Yu, En, et al.
Pubblicazione: (2023)
di: Yu, En, et al.
Pubblicazione: (2023)
OneChart: Purify the Chart Structural Extraction via One Auxiliary Token
di: Chen, Jinyue, et al.
Pubblicazione: (2024)
di: Chen, Jinyue, et al.
Pubblicazione: (2024)
DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation
di: Peng, Yuang, et al.
Pubblicazione: (2024)
di: Peng, Yuang, et al.
Pubblicazione: (2024)
Focus Anywhere for Fine-grained Multi-page Document Understanding
di: Liu, Chenglong, et al.
Pubblicazione: (2024)
di: Liu, Chenglong, et al.
Pubblicazione: (2024)
Taming Teacher Forcing for Masked Autoregressive Video Generation
di: Zhou, Deyu, et al.
Pubblicazione: (2025)
di: Zhou, Deyu, et al.
Pubblicazione: (2025)
ShapeLLM: Universal 3D Object Understanding for Embodied Interaction
di: Qi, Zekun, et al.
Pubblicazione: (2024)
di: Qi, Zekun, et al.
Pubblicazione: (2024)
Positional Prompt Tuning for Efficient 3D Representation Learning
di: Zhang, Shaochen, et al.
Pubblicazione: (2024)
di: Zhang, Shaochen, et al.
Pubblicazione: (2024)
SpatialEvo: Self-Evolving Spatial Intelligence via Deterministic Geometric Environments
di: Li, Dinging, et al.
Pubblicazione: (2026)
di: Li, Dinging, et al.
Pubblicazione: (2026)
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
di: NextStep Team, et al.
Pubblicazione: (2025)
di: NextStep Team, et al.
Pubblicazione: (2025)
Learning Humanoid End-Effector Control for Open-Vocabulary Visual Loco-Manipulation
di: Dong, Runpei, et al.
Pubblicazione: (2026)
di: Dong, Runpei, et al.
Pubblicazione: (2026)
Thermo-VL: Extending Vision-Language Models to Thermal Infrared Perception
di: Thushara, Rusiru, et al.
Pubblicazione: (2026)
di: Thushara, Rusiru, et al.
Pubblicazione: (2026)
LCV2I: Communication-Efficient and High-Performance Collaborative Perception Framework with Low-Resolution LiDAR
di: Feng, Xinxin, et al.
Pubblicazione: (2025)
di: Feng, Xinxin, et al.
Pubblicazione: (2025)
ULTRA: Unified Multimodal Control for Autonomous Humanoid Whole-Body Loco-Manipulation
di: He, Xialin, et al.
Pubblicazione: (2026)
di: He, Xialin, et al.
Pubblicazione: (2026)
Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
di: Liu, Ruyang, et al.
Pubblicazione: (2025)
di: Liu, Ruyang, et al.
Pubblicazione: (2025)
From Web to Pixels: Bringing Agentic Search into Visual Perception
di: Yang, Bokang, et al.
Pubblicazione: (2026)
di: Yang, Bokang, et al.
Pubblicazione: (2026)
AWRaCLe: All-Weather Image Restoration using Visual In-Context Learning
di: Rajagopalan, Sudarshan, et al.
Pubblicazione: (2024)
di: Rajagopalan, Sudarshan, et al.
Pubblicazione: (2024)
Not All Tokens Need 40 Steps: Heterogeneous Step Allocation in Diffusion Transformers for Efficient Video Generation
di: Chu, Ernie, et al.
Pubblicazione: (2026)
di: Chu, Ernie, et al.
Pubblicazione: (2026)
Low-rank Adaptation-based All-Weather Removal for Autonomous Navigation
di: Rajagopalan, Sudarshan, et al.
Pubblicazione: (2024)
di: Rajagopalan, Sudarshan, et al.
Pubblicazione: (2024)
MedCL: Learning Consistent Anatomy Distribution for Scribble-supervised Medical Image Segmentation
di: Zhang, Ke, et al.
Pubblicazione: (2025)
di: Zhang, Ke, et al.
Pubblicazione: (2025)
RemoteVAR: Autoregressive Visual Modeling for Remote Sensing Change Detection
di: Korkmaz, Yilmaz, et al.
Pubblicazione: (2026)
di: Korkmaz, Yilmaz, et al.
Pubblicazione: (2026)
Hyp-OC: Hyperbolic One Class Classification for Face Anti-Spoofing
di: Narayan, Kartik, et al.
Pubblicazione: (2024)
di: Narayan, Kartik, et al.
Pubblicazione: (2024)
ModelMix: A New Model-Mixup Strategy to Minimize Vicinal Risk across Tasks for Few-scribble based Cardiac Segmentation
di: Zhang, Ke, et al.
Pubblicazione: (2024)
di: Zhang, Ke, et al.
Pubblicazione: (2024)
Active Learning for Vision-Language Models
di: Safaei, Bardia, et al.
Pubblicazione: (2024)
di: Safaei, Bardia, et al.
Pubblicazione: (2024)
Implicit Neural Representations: A Signal Processing Perspective
di: Jayasundara, Dhananjaya, et al.
Pubblicazione: (2026)
di: Jayasundara, Dhananjaya, et al.
Pubblicazione: (2026)
Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation
di: Jiang, Haichao, et al.
Pubblicazione: (2026)
di: Jiang, Haichao, et al.
Pubblicazione: (2026)
DeepSeek-OCR: Contexts Optical Compression
di: Wei, Haoran, et al.
Pubblicazione: (2025)
di: Wei, Haoran, et al.
Pubblicazione: (2025)
DeepSeek-OCR 2: Visual Causal Flow
di: Wei, Haoran, et al.
Pubblicazione: (2026)
di: Wei, Haoran, et al.
Pubblicazione: (2026)
On Asymmetric Optimization of Reasoning and Perception in Vision-Language Model Post-Training
di: Wu, Xueqing, et al.
Pubblicazione: (2026)
di: Wu, Xueqing, et al.
Pubblicazione: (2026)
Segment Anything, Even Occluded
di: Tai, Wei-En, et al.
Pubblicazione: (2025)
di: Tai, Wei-En, et al.
Pubblicazione: (2025)
R2SM: Referring and Reasoning for Selective Masks
di: Shih, Yu-Lin, et al.
Pubblicazione: (2025)
di: Shih, Yu-Lin, et al.
Pubblicazione: (2025)
ViSA-Enhanced Aerial VLN: A Visual-Spatial Reasoning Enhanced Framework for Aerial Vision-Language Navigation
di: Tong, Haoyu, et al.
Pubblicazione: (2026)
di: Tong, Haoyu, et al.
Pubblicazione: (2026)
Dreamguider: Improved Training free Diffusion-based Conditional Generation
di: Nair, Nithin Gopalakrishnan, et al.
Pubblicazione: (2024)
di: Nair, Nithin Gopalakrishnan, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Perception-R1: Pioneering Perception Policy with Reinforcement Learning
di: Yu, En, et al.
Pubblicazione: (2025) -
Unhackable Temporal Rewarding for Scalable Video MLLMs
di: Yu, En, et al.
Pubblicazione: (2025) -
Exploring Recurrent Long-term Temporal Fusion for Multi-view 3D Perception
di: Han, Chunrui, et al.
Pubblicazione: (2023) -
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
di: Wei, Yana, et al.
Pubblicazione: (2025) -
Small Language Model Meets with Reinforced Vision Vocabulary
di: Wei, Haoran, et al.
Pubblicazione: (2024)