LoopViT: Scaling Visual ARC with Looped Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Shu, Wen-Jie, Qiu, Xuerui, Zhu, Rui-Jie, Chen, Harold Haodong, Liu, Yexin, Yang, Harry |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
Temporal Regularization Makes Your Video Generator Stronger
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
ViPO: Visual Preference Optimization at Scale
von: Li, Ming, et al.
Veröffentlicht: (2026)
von: Li, Ming, et al.
Veröffentlicht: (2026)
LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization
von: Wu, Xianfeng, et al.
Veröffentlicht: (2025)
von: Wu, Xianfeng, et al.
Veröffentlicht: (2025)
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
von: Liu, Yexin, et al.
Veröffentlicht: (2025)
von: Liu, Yexin, et al.
Veröffentlicht: (2025)
TiViBench: Benchmarking Think-in-Video Reasoning for Video Generative Models
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
ELT: Elastic Looped Transformers for Visual Generation
von: Goyal, Sahil, et al.
Veröffentlicht: (2026)
von: Goyal, Sahil, et al.
Veröffentlicht: (2026)
Quantized Spike-driven Transformer
von: Qiu, Xuerui, et al.
Veröffentlicht: (2025)
von: Qiu, Xuerui, et al.
Veröffentlicht: (2025)
Retina Vision Transformer (RetinaViT): Introducing Scaled Patches into Vision Transformers
von: Shu, Yuyang, et al.
Veröffentlicht: (2024)
von: Shu, Yuyang, et al.
Veröffentlicht: (2024)
LoopSparseGS: Loop Based Sparse-View Friendly Gaussian Splatting
von: Bao, Zhenyu, et al.
Veröffentlicht: (2024)
von: Bao, Zhenyu, et al.
Veröffentlicht: (2024)
AutoLayout: Closed-Loop Layout Synthesis via Slow-Fast Collaborative Reasoning
von: Chen, Weixing, et al.
Veröffentlicht: (2025)
von: Chen, Weixing, et al.
Veröffentlicht: (2025)
V-Loop: Visual Logical Loop Verification for Hallucination Detection in Medical Visual Question Answering
von: Jin, Mengyuan, et al.
Veröffentlicht: (2026)
von: Jin, Mengyuan, et al.
Veröffentlicht: (2026)
Beyond Generation: Unlocking Universal Editing via Self-Supervised Fine-Tuning
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2024)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2024)
ViEEG: Hierarchical Visual Neural Representation for EEG Brain Decoding
von: Liu, Minxu, et al.
Veröffentlicht: (2025)
von: Liu, Minxu, et al.
Veröffentlicht: (2025)
Hierarchical Fine-grained Preference Optimization for Physically Plausible Video Generation
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025)
Think Visually, Reason Textually: Vision-Language Synergy in ARC
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
von: Zhang, Beichen, et al.
Veröffentlicht: (2025)
UniViTAR: Unified Vision Transformer with Native Resolution
von: Qiao, Limeng, et al.
Veröffentlicht: (2025)
von: Qiao, Limeng, et al.
Veröffentlicht: (2025)
LoopDB: A Loop Closure Dataset for Large Scale Simultaneous Localization and Mapping
von: Nakshbandi, Mohammad-Maher, et al.
Veröffentlicht: (2025)
von: Nakshbandi, Mohammad-Maher, et al.
Veröffentlicht: (2025)
CLOVA: A Closed-Loop Visual Assistant with Tool Usage and Update
von: Gao, Zhi, et al.
Veröffentlicht: (2023)
von: Gao, Zhi, et al.
Veröffentlicht: (2023)
User-in-the-Loop View Sampling with Error Peaking Visualization
von: Yasunaga, Ayaka, et al.
Veröffentlicht: (2025)
von: Yasunaga, Ayaka, et al.
Veröffentlicht: (2025)
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2026)
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2026)
LoopSplat: Loop Closure by Registering 3D Gaussian Splats
von: Zhu, Liyuan, et al.
Veröffentlicht: (2024)
von: Zhu, Liyuan, et al.
Veröffentlicht: (2024)
DD-VNB: A Depth-based Dual-Loop Framework for Real-time Visually Navigated Bronchoscopy
von: Tian, Qingyao, et al.
Veröffentlicht: (2024)
von: Tian, Qingyao, et al.
Veröffentlicht: (2024)
LORTSAR: Low-Rank Transformer for Skeleton-based Action Recognition
von: Oraki, Soroush, et al.
Veröffentlicht: (2024)
von: Oraki, Soroush, et al.
Veröffentlicht: (2024)
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
Human-in-the-Loop Visual Re-ID for Population Size Estimation
von: Perez, Gustavo, et al.
Veröffentlicht: (2023)
von: Perez, Gustavo, et al.
Veröffentlicht: (2023)
Towards Effective Human-in-the-Loop Assistive AI Agents
von: Bellos, Filippos, et al.
Veröffentlicht: (2025)
von: Bellos, Filippos, et al.
Veröffentlicht: (2025)
Render-in-the-Loop: Vector Graphics Generation via Visual Self-Feedback
von: Liang, Guotao, et al.
Veröffentlicht: (2026)
von: Liang, Guotao, et al.
Veröffentlicht: (2026)
Déjà View: Looping Transformers for Multi-View 3D Reconstruction
von: Burzio, Alessandro, et al.
Veröffentlicht: (2026)
von: Burzio, Alessandro, et al.
Veröffentlicht: (2026)
Exploiting Lightweight Hierarchical ViT and Dynamic Framework for Efficient Visual Tracking
von: Kang, Ben, et al.
Veröffentlicht: (2025)
von: Kang, Ben, et al.
Veröffentlicht: (2025)
NaViL: Rethinking Scaling Properties of Native Multimodal Large Language Models under Data Constraints
von: Tian, Changyao, et al.
Veröffentlicht: (2025)
von: Tian, Changyao, et al.
Veröffentlicht: (2025)
Scaling Spike-driven Transformer with Efficient Spike Firing Approximation Training
von: Yao, Man, et al.
Veröffentlicht: (2024)
von: Yao, Man, et al.
Veröffentlicht: (2024)
LoopNet: A Multitasking Few-Shot Learning Approach for Loop Closure in Large Scale SLAM
von: Nakshbandi, Mohammad-Maher, et al.
Veröffentlicht: (2025)
von: Nakshbandi, Mohammad-Maher, et al.
Veröffentlicht: (2025)
Autonomous Imagination: Closed-Loop Decomposition of Visual-to-Textual Conversion in Visual Reasoning for Multimodal Large Language Models
von: Liu, Jingming, et al.
Veröffentlicht: (2024)
von: Liu, Jingming, et al.
Veröffentlicht: (2024)
ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?
von: Han, Haonan, et al.
Veröffentlicht: (2026)
von: Han, Haonan, et al.
Veröffentlicht: (2026)
EA-ViT: Efficient Adaptation for Elastic Vision Transformer
von: Zhu, Chen, et al.
Veröffentlicht: (2025)
von: Zhu, Chen, et al.
Veröffentlicht: (2025)
Visual Loop Closure Detection Through Deep Graph Consensus
von: Büchner, Martin, et al.
Veröffentlicht: (2025)
von: Büchner, Martin, et al.
Veröffentlicht: (2025)
Interactive Garment Recommendation with User in the Loop
von: Becattini, Federico, et al.
Veröffentlicht: (2024)
von: Becattini, Federico, et al.
Veröffentlicht: (2024)
LA-Sign: Looped Transformers with Geometry-aware Alignment for Skeleton-based Sign Language Recognition
von: Pu, Muxin, et al.
Veröffentlicht: (2026)
von: Pu, Muxin, et al.
Veröffentlicht: (2026)
TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis
von: Chen, Shunian, et al.
Veröffentlicht: (2025)
von: Chen, Shunian, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Go with Your Gut: Scaling Confidence for Autoregressive Image Generation
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025) -
Temporal Regularization Makes Your Video Generator Stronger
von: Chen, Harold Haodong, et al.
Veröffentlicht: (2025) -
ViPO: Visual Preference Optimization at Scale
von: Li, Ming, et al.
Veröffentlicht: (2026) -
LightGen: Efficient Image Generation through Knowledge Distillation and Direct Preference Optimization
von: Wu, Xianfeng, et al.
Veröffentlicht: (2025) -
AlignVid: Training-Free Attention Scaling for Semantic Fidelity in Text-Guided Image-to-Video Generation
von: Liu, Yexin, et al.
Veröffentlicht: (2025)