FreeAction: Training-Free Techniques for Enhanced Fidelity of Trajectory-to-Video Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Seungwook, Lee, Seunghyeon, Cho, Minsu |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Harnessing the Power of Training-Free Techniques in Text-to-2D Generation for Text-to-3D Generation via Score Distillation Sampling
by: Lee, Junhong, et al.
Published: (2025)
by: Lee, Junhong, et al.
Published: (2025)
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
by: Kim, Seungwook, et al.
Published: (2026)
by: Kim, Seungwook, et al.
Published: (2026)
Similarity-Aware Selective State-Space Modeling for Semantic Correspondence
by: Kim, Seungwook, et al.
Published: (2025)
by: Kim, Seungwook, et al.
Published: (2025)
SpaCeFormer: Fast Proposal-Free Open-Vocabulary 3D Instance Segmentation
by: Choy, Chris, et al.
Published: (2026)
by: Choy, Chris, et al.
Published: (2026)
Affostruction: 3D Affordance Grounding with Generative Reconstruction
by: Park, Chunghyun, et al.
Published: (2026)
by: Park, Chunghyun, et al.
Published: (2026)
DextER: Language-driven Dexterous Grasp Generation with Embodied Reasoning
by: Lee, Junha, et al.
Published: (2026)
by: Lee, Junha, et al.
Published: (2026)
CorrespondentDream: Enhancing 3D Fidelity of Text-to-3D using Cross-View Correspondences
by: Kim, Seungwook, et al.
Published: (2024)
by: Kim, Seungwook, et al.
Published: (2024)
Learning SO(3)-Invariant Semantic Correspondence via Local Shape Transform
by: Park, Chunghyun, et al.
Published: (2024)
by: Park, Chunghyun, et al.
Published: (2024)
Multi-view Image Prompted Multi-view Diffusion for Improved 3D Generation
by: Kim, Seungwook, et al.
Published: (2024)
by: Kim, Seungwook, et al.
Published: (2024)
FreeOcc: Training-Free Embodied Open-Vocabulary Occupancy Prediction
by: Jiang, Zeyu, et al.
Published: (2026)
by: Jiang, Zeyu, et al.
Published: (2026)
Autoregressive Meta-Actions for Unified Controllable Trajectory Generation
by: Zhao, Jianbo, et al.
Published: (2025)
by: Zhao, Jianbo, et al.
Published: (2025)
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
by: Kim, Dongwon, et al.
Published: (2026)
by: Kim, Dongwon, et al.
Published: (2026)
3D Equivariant Pose Regression via Direct Wigner-D Harmonics Prediction
by: Lee, Jongmin, et al.
Published: (2024)
by: Lee, Jongmin, et al.
Published: (2024)
RapidMV: Leveraging Spatio-Angular Representations for Efficient and Consistent Text-to-Multi-View Synthesis
by: Kim, Seungwook, et al.
Published: (2025)
by: Kim, Seungwook, et al.
Published: (2025)
Closed-Loop Action Chunks with Dynamic Corrections for Training-Free Diffusion Policy
by: Wu, Pengyuan, et al.
Published: (2026)
by: Wu, Pengyuan, et al.
Published: (2026)
Precise Action-to-Video Generation Through Visual Action Prompts
by: Wang, Yuang, et al.
Published: (2025)
by: Wang, Yuang, et al.
Published: (2025)
Video Generation with Learned Action Prior
by: Sarkar, Meenakshi, et al.
Published: (2024)
by: Sarkar, Meenakshi, et al.
Published: (2024)
VLA-InfoEntropy: A Training-Free Vision-Attention Information Entropy Approach for Vision-Language-Action Models Inference Acceleration and Success
by: Liu, Chuhang, et al.
Published: (2026)
by: Liu, Chuhang, et al.
Published: (2026)
From Generated Human Videos to Physically Plausible Robot Trajectories
by: Ni, James, et al.
Published: (2025)
by: Ni, James, et al.
Published: (2025)
UniSkill: Imitating Human Videos via Cross-Embodiment Skill Representations
by: Kim, Hanjung, et al.
Published: (2025)
by: Kim, Hanjung, et al.
Published: (2025)
LIVE-GS: Online LiDAR-Inertial-Visual State Estimation and Globally Consistent Mapping with 3D Gaussian Splatting
by: Park, Jaeseok, et al.
Published: (2025)
by: Park, Jaeseok, et al.
Published: (2025)
GeneralVLA: Generalizable Vision-Language-Action Models with Knowledge-Guided Trajectory Planning
by: Ma, Guoqing, et al.
Published: (2026)
by: Ma, Guoqing, et al.
Published: (2026)
Post-Training and Test-Time Scaling of Generative Agent Behavior Models for Interactive Autonomous Driving
by: Seong, Hyunki, et al.
Published: (2025)
by: Seong, Hyunki, et al.
Published: (2025)
Mask2IV: Interaction-Centric Video Generation via Mask Trajectories
by: Li, Gen, et al.
Published: (2025)
by: Li, Gen, et al.
Published: (2025)
VAG: Dual-Stream Video-Action Generation for Embodied Data Synthesis
by: Lang, Xiaolei, et al.
Published: (2026)
by: Lang, Xiaolei, et al.
Published: (2026)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
by: Kim, Manjin, et al.
Published: (2026)
by: Kim, Manjin, et al.
Published: (2026)
Unified Video Action Model
by: Li, Shuang, et al.
Published: (2025)
by: Li, Shuang, et al.
Published: (2025)
GAD-Generative Learning for HD Map-Free Autonomous Driving
by: Sun, Weijian, et al.
Published: (2024)
by: Sun, Weijian, et al.
Published: (2024)
UniGround: Universal 3D Visual Grounding via Training-Free Scene Parsing
by: Zhang, Jiaxi, et al.
Published: (2026)
by: Zhang, Jiaxi, et al.
Published: (2026)
Action Images: End-to-End Policy Learning via Multiview Video Generation
by: Zhen, Haoyu, et al.
Published: (2026)
by: Zhen, Haoyu, et al.
Published: (2026)
LAD-Drive: Bridging Language and Trajectory with Action-Aware Diffusion Transformers
by: Schmidt, Fabian, et al.
Published: (2026)
by: Schmidt, Fabian, et al.
Published: (2026)
Classification Matters: Improving Video Action Detection with Class-Specific Attention
by: Lee, Jinsung, et al.
Published: (2024)
by: Lee, Jinsung, et al.
Published: (2024)
HEAT: Heterogeneous End-to-End Autonomous Driving via Trajectory-Guided World Models
by: Cho, Hoonhee, et al.
Published: (2026)
by: Cho, Hoonhee, et al.
Published: (2026)
3D Geometric Shape Assembly via Efficient Point Cloud Matching
by: Lee, Nahyuk, et al.
Published: (2024)
by: Lee, Nahyuk, et al.
Published: (2024)
RoboCurate: Harnessing Diversity with Action-Verified Neural Trajectory for Robot Learning
by: Kim, Seungku, et al.
Published: (2026)
by: Kim, Seungku, et al.
Published: (2026)
Scalable Benchmarking and Robust Learning for Noise-Free Ego-Motion and 3D Reconstruction from Noisy Video
by: Xu, Xiaohao, et al.
Published: (2025)
by: Xu, Xiaohao, et al.
Published: (2025)
Thermal Chameleon: Task-Adaptive Tone-mapping for Radiometric Thermal-Infrared images
by: Lee, Dong-Guw, et al.
Published: (2024)
by: Lee, Dong-Guw, et al.
Published: (2024)
Foundation Feature-Driven Online End-Effector Pose Estimation: A Marker-Free and Learning-Free Approach
by: Wu, Tianshu, et al.
Published: (2025)
by: Wu, Tianshu, et al.
Published: (2025)
PosA-VLA: Enhancing Action Generation via Pose-Conditioned Anchor Attention
by: Li, Ziwen, et al.
Published: (2025)
by: Li, Ziwen, et al.
Published: (2025)
What's Wrong with the Absolute Trajectory Error?
by: Lee, Seong Hun, et al.
Published: (2022)
by: Lee, Seong Hun, et al.
Published: (2022)
Similar Items
-
Harnessing the Power of Training-Free Techniques in Text-to-2D Generation for Text-to-3D Generation via Score Distillation Sampling
by: Lee, Junhong, et al.
Published: (2025) -
Improving Text-to-Image Generation with Intrinsic Self-Confidence Rewards
by: Kim, Seungwook, et al.
Published: (2026) -
Similarity-Aware Selective State-Space Modeling for Semantic Correspondence
by: Kim, Seungwook, et al.
Published: (2025) -
SpaCeFormer: Fast Proposal-Free Open-Vocabulary 3D Instance Segmentation
by: Choy, Chris, et al.
Published: (2026) -
Affostruction: 3D Affordance Grounding with Generative Reconstruction
by: Park, Chunghyun, et al.
Published: (2026)