Token Warping Helps MLLMs Look from Nearby Viewpoints
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Phillip Y., Park, Chanho, Park, Mingue, Yoo, Seungwoo, Koo, Juil, Sung, Minhyuk |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Posterior Distillation Sampling
von: Koo, Juil, et al.
Veröffentlicht: (2023)
von: Koo, Juil, et al.
Veröffentlicht: (2023)
Neural Pose Representation Learning for Generating and Transferring Non-Rigid Object Poses
von: Yoo, Seungwoo, et al.
Veröffentlicht: (2024)
von: Yoo, Seungwoo, et al.
Veröffentlicht: (2024)
SALAD: Part-Level Latent Diffusion for 3D Shape Generation and Manipulation
von: Koo, Juil, et al.
Veröffentlicht: (2023)
von: Koo, Juil, et al.
Veröffentlicht: (2023)
BoxSplitGen: A Generative Model for 3D Part Bounding Boxes in Varying Granularity
von: Koo, Juil, et al.
Veröffentlicht: (2026)
von: Koo, Juil, et al.
Veröffentlicht: (2026)
DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
von: Park, Mingue, et al.
Veröffentlicht: (2025)
von: Park, Mingue, et al.
Veröffentlicht: (2025)
Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection
von: Koo, Juil, et al.
Veröffentlicht: (2025)
von: Koo, Juil, et al.
Veröffentlicht: (2025)
SyncTweedies: A General Generative Framework Based on Synchronized Diffusions
von: Kim, Jaihoon, et al.
Veröffentlicht: (2024)
von: Kim, Jaihoon, et al.
Veröffentlicht: (2024)
Proxy-Free Gaussian Splats Deformation with Splat-Based Surface Estimation
von: Kim, Jaeyeong, et al.
Veröffentlicht: (2025)
von: Kim, Jaeyeong, et al.
Veröffentlicht: (2025)
Perspective-Aware Reasoning in Vision-Language Models via Mental Imagery Simulation
von: Lee, Phillip Y., et al.
Veröffentlicht: (2025)
von: Lee, Phillip Y., et al.
Veröffentlicht: (2025)
ReGround: Improving Textual and Spatial Grounding at No Cost
von: Lee, Phillip Y., et al.
Veröffentlicht: (2024)
von: Lee, Phillip Y., et al.
Veröffentlicht: (2024)
VideoHandles: Editing 3D Object Compositions in Videos Using Video Generative Priors
von: Koo, Juil, et al.
Veröffentlicht: (2025)
von: Koo, Juil, et al.
Veröffentlicht: (2025)
GrounDiT: Grounding Diffusion Transformers via Noisy Patch Transplantation
von: Lee, Phillip Y., et al.
Veröffentlicht: (2024)
von: Lee, Phillip Y., et al.
Veröffentlicht: (2024)
As-Plausible-As-Possible: Plausibility-Aware Mesh Deformation Using 2D Diffusion Priors
von: Yoo, Seungwoo, et al.
Veröffentlicht: (2023)
von: Yoo, Seungwoo, et al.
Veröffentlicht: (2023)
Unconditional Priors Matter! Improving Conditional Generation of Fine-Tuned Diffusion Models
von: Phunyaphibarn, Prin, et al.
Veröffentlicht: (2025)
von: Phunyaphibarn, Prin, et al.
Veröffentlicht: (2025)
Drifting Field Policy: A One-Step Generative Policy via Wasserstein Gradient Flow
von: Koo, Juil, et al.
Veröffentlicht: (2026)
von: Koo, Juil, et al.
Veröffentlicht: (2026)
Domain Generalization for Person Re-identification: A Survey Towards Domain-Agnostic Person Matching
von: Lee, Hyeonseo, et al.
Veröffentlicht: (2025)
von: Lee, Hyeonseo, et al.
Veröffentlicht: (2025)
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding
von: Kim, Jiwan, et al.
Veröffentlicht: (2026)
von: Kim, Jiwan, et al.
Veröffentlicht: (2026)
PartSTAD: 2D-to-3D Part Segmentation Task Adaptation
von: Kim, Hyunjin, et al.
Veröffentlicht: (2024)
von: Kim, Hyunjin, et al.
Veröffentlicht: (2024)
DiffusionRollout: Uncertainty-Aware Rollout Planning in Long-Horizon PDE Solving
von: Yoo, Seungwoo, et al.
Veröffentlicht: (2026)
von: Yoo, Seungwoo, et al.
Veröffentlicht: (2026)
BézierFlow: Learning Bézier Stochastic Interpolant Schedulers for Few-Step Generation
von: Min, Yunhong, et al.
Veröffentlicht: (2025)
von: Min, Yunhong, et al.
Veröffentlicht: (2025)
GridPrune: From "Where to Look" to "What to Select" in Visual Token Pruning for MLLMs
von: Duan, Yuxiang, et al.
Veröffentlicht: (2025)
von: Duan, Yuxiang, et al.
Veröffentlicht: (2025)
MorphGS: Morphology-Adaptive Articulated 3D Motion Transfer from Videos
von: Kim, Taeyeon, et al.
Veröffentlicht: (2026)
von: Kim, Taeyeon, et al.
Veröffentlicht: (2026)
Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment
von: Park, Jonghyun, et al.
Veröffentlicht: (2025)
von: Park, Jonghyun, et al.
Veröffentlicht: (2025)
Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
von: Dalal, Dwip, et al.
Veröffentlicht: (2025)
von: Dalal, Dwip, et al.
Veröffentlicht: (2025)
PairFlow: Closed-Form Source-Target Coupling for Few-Step Generation in Discrete Flow Models
von: Park, Mingue, et al.
Veröffentlicht: (2025)
von: Park, Mingue, et al.
Veröffentlicht: (2025)
Zero-Shot Video Translation via Token Warping
von: Zhu, Haiming, et al.
Veröffentlicht: (2024)
von: Zhu, Haiming, et al.
Veröffentlicht: (2024)
vid-TLDR: Training Free Token merging for Light-weight Video Transformer
von: Choi, Joonmyung, et al.
Veröffentlicht: (2024)
von: Choi, Joonmyung, et al.
Veröffentlicht: (2024)
Occupancy-Based Dual Contouring
von: Hwang, Jisung, et al.
Veröffentlicht: (2024)
von: Hwang, Jisung, et al.
Veröffentlicht: (2024)
Coherent Human-Scene Reconstruction from Multi-Person Multi-View Video in a Single Pass
von: Kim, Sangmin, et al.
Veröffentlicht: (2026)
von: Kim, Sangmin, et al.
Veröffentlicht: (2026)
Extend3D: Town-Scale 3D Generation
von: Yoon, Seungwoo, et al.
Veröffentlicht: (2026)
von: Yoon, Seungwoo, et al.
Veröffentlicht: (2026)
MemBench: Memorized Image Trigger Prompt Dataset for Diffusion Models
von: Hong, Chunsan, et al.
Veröffentlicht: (2024)
von: Hong, Chunsan, et al.
Veröffentlicht: (2024)
Customizing Text-to-Image Diffusion with Object Viewpoint Control
von: Kumari, Nupur, et al.
Veröffentlicht: (2024)
von: Kumari, Nupur, et al.
Veröffentlicht: (2024)
MoCHA-former: Moiré-Conditioned Hybrid Adaptive Transformer for Video Demoiréing
von: Sung, Jeahun, et al.
Veröffentlicht: (2025)
von: Sung, Jeahun, et al.
Veröffentlicht: (2025)
Memory Helps, but Confabulation Misleads: Understanding Streaming Events in Videos with MLLMs
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2025)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2025)
ARC-NeRF: Area Ray Casting for Broader Unseen View Coverage in Few-shot Object Rendering
von: Seo, Seunghyeon, et al.
Veröffentlicht: (2024)
von: Seo, Seunghyeon, et al.
Veröffentlicht: (2024)
MatLat: Material Latent Space for PBR Texture Generation
von: Yeo, Kyeongmin, et al.
Veröffentlicht: (2025)
von: Yeo, Kyeongmin, et al.
Veröffentlicht: (2025)
InterHandGen: Two-Hand Interaction Generation via Cascaded Reverse Diffusion
von: Lee, Jihyun, et al.
Veröffentlicht: (2024)
von: Lee, Jihyun, et al.
Veröffentlicht: (2024)
Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens
von: Lu, Xinxuan, et al.
Veröffentlicht: (2026)
von: Lu, Xinxuan, et al.
Veröffentlicht: (2026)
StochSync: Stochastic Diffusion Synchronization for Image Generation in Arbitrary Spaces
von: Yeo, Kyeongmin, et al.
Veröffentlicht: (2025)
von: Yeo, Kyeongmin, et al.
Veröffentlicht: (2025)
LookingGlass: Generative Anamorphoses via Laplacian Pyramid Warping
von: Chang, Pascal, et al.
Veröffentlicht: (2025)
von: Chang, Pascal, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Posterior Distillation Sampling
von: Koo, Juil, et al.
Veröffentlicht: (2023) -
Neural Pose Representation Learning for Generating and Transferring Non-Rigid Object Poses
von: Yoo, Seungwoo, et al.
Veröffentlicht: (2024) -
SALAD: Part-Level Latent Diffusion for 3D Shape Generation and Manipulation
von: Koo, Juil, et al.
Veröffentlicht: (2023) -
BoxSplitGen: A Generative Model for 3D Part Bounding Boxes in Varying Granularity
von: Koo, Juil, et al.
Veröffentlicht: (2026) -
DiverseVAR: Balancing Diversity and Quality of Next-Scale Visual Autoregressive Models
von: Park, Mingue, et al.
Veröffentlicht: (2025)