PALM: Predicting Actions through Language Models
Fuente:
arXiv
Saved in:
| Main Authors: | Kim, Sanghwan, Huang, Daoji, Xian, Yongqin, Hilliges, Otmar, Van Gool, Luc, Wang, Xi |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation
by: Pasca, Razvan-George, et al.
Published: (2023)
by: Pasca, Razvan-George, et al.
Published: (2023)
Leveraging Driver Field-of-View for Multimodal Ego-Trajectory Prediction
by: Akbiyik, M. Eren, et al.
Published: (2023)
by: Akbiyik, M. Eren, et al.
Published: (2023)
EgoGaussian: Dynamic Scene Understanding from Egocentric Video with 3D Gaussian Splatting
by: Zhang, Daiwei, et al.
Published: (2024)
by: Zhang, Daiwei, et al.
Published: (2024)
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
by: Kim, Sanghwan, et al.
Published: (2025)
by: Kim, Sanghwan, et al.
Published: (2025)
MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
by: Segu, Mattia, et al.
Published: (2025)
by: Segu, Mattia, et al.
Published: (2025)
FINER: MLLMs Hallucinate under Fine-grained Negative Queries
by: Xiao, Rui, et al.
Published: (2026)
by: Xiao, Rui, et al.
Published: (2026)
CADS: Unleashing the Diversity of Diffusion Models through Condition-Annealed Sampling
by: Sadat, Seyedmorteza, et al.
Published: (2023)
by: Sadat, Seyedmorteza, et al.
Published: (2023)
SiTH: Single-view Textured Human Reconstruction with Image-Conditioned Diffusion
by: Ho, Hsuan-I, et al.
Published: (2023)
by: Ho, Hsuan-I, et al.
Published: (2023)
Eliminating Oversaturation and Artifacts of High Guidance Scales in Diffusion Models
by: Sadat, Seyedmorteza, et al.
Published: (2024)
by: Sadat, Seyedmorteza, et al.
Published: (2024)
AvatarPose: Avatar-guided 3D Pose Estimation of Close Human Interaction from Sparse Multi-view Videos
by: Lu, Feichi, et al.
Published: (2024)
by: Lu, Feichi, et al.
Published: (2024)
Optimizing against Infeasible Inclusions from Data for Semantic Segmentation through Morphology
by: Basu, Shamik, et al.
Published: (2024)
by: Basu, Shamik, et al.
Published: (2024)
No Training, No Problem: Rethinking Classifier-Free Guidance for Diffusion Models
by: Sadat, Seyedmorteza, et al.
Published: (2024)
by: Sadat, Seyedmorteza, et al.
Published: (2024)
GaussianVLM: Scene-centric 3D Vision-Language Models using Language-aligned Gaussian Splats for Embodied Reasoning and Beyond
by: Halacheva, Anna-Maria, et al.
Published: (2025)
by: Halacheva, Anna-Maria, et al.
Published: (2025)
WANDR: Intention-guided Human Motion Generation
by: Diomataris, Markos, et al.
Published: (2024)
by: Diomataris, Markos, et al.
Published: (2024)
GraspXL: Generating Grasping Motions for Diverse Objects at Scale
by: Zhang, Hui, et al.
Published: (2024)
by: Zhang, Hui, et al.
Published: (2024)
LiteVAE: Lightweight and Efficient Variational Autoencoders for Latent Diffusion Models
by: Sadat, Seyedmorteza, et al.
Published: (2024)
by: Sadat, Seyedmorteza, et al.
Published: (2024)
LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation
by: Miao, Yang, et al.
Published: (2025)
by: Miao, Yang, et al.
Published: (2025)
Distilling ODE Solvers of Diffusion Models into Smaller Steps
by: Kim, Sanghwan, et al.
Published: (2023)
by: Kim, Sanghwan, et al.
Published: (2023)
StateSpaceDiffuser: Bringing Long Context to Diffusion World Models
by: Savov, Nedko, et al.
Published: (2025)
by: Savov, Nedko, et al.
Published: (2025)
Test-time Training for Hyperspectral Image Super-resolution
by: Li, Ke, et al.
Published: (2024)
by: Li, Ke, et al.
Published: (2024)
Bayesian Self-Training for Semi-Supervised 3D Segmentation
by: Unal, Ozan, et al.
Published: (2024)
by: Unal, Ozan, et al.
Published: (2024)
TRAVL: A Recipe for Making Video-Language Models Better Judges of Physics Implausibility
by: Motamed, Saman, et al.
Published: (2025)
by: Motamed, Saman, et al.
Published: (2025)
Graph Transformer GANs with Graph Masked Modeling for Architectural Layout Generation
by: Tang, Hao, et al.
Published: (2024)
by: Tang, Hao, et al.
Published: (2024)
MatIR: A Hybrid Mamba-Transformer Image Restoration Model
by: Wen, Juan, et al.
Published: (2025)
by: Wen, Juan, et al.
Published: (2025)
Investigating the Effectiveness of Cross-Attention to Unlock Zero-Shot Editing of Text-to-Video Diffusion Models
by: Motamed, Saman, et al.
Published: (2024)
by: Motamed, Saman, et al.
Published: (2024)
Learning to Prompt with Text Only Supervision for Vision-Language Models
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
by: Khattak, Muhammad Uzair, et al.
Published: (2024)
Camera-Only 3D Panoptic Scene Completion for Autonomous Driving through Differentiable Object Shapes
by: Marinello, Nicola, et al.
Published: (2025)
by: Marinello, Nicola, et al.
Published: (2025)
Occam's LGS: An Efficient Approach for Language Gaussian Splatting
by: Cheng, Jiahuan, et al.
Published: (2024)
by: Cheng, Jiahuan, et al.
Published: (2024)
SIGHT: Synthesizing Image-Text Conditioned and Geometry-Guided 3D Hand-Object Trajectories
by: Gavryushin, Alexey, et al.
Published: (2025)
by: Gavryushin, Alexey, et al.
Published: (2025)
Human Hair Reconstruction with Strand-Aligned 3D Gaussians
by: Zakharov, Egor, et al.
Published: (2024)
by: Zakharov, Egor, et al.
Published: (2024)
EvenNICER-SLAM: Event-based Neural Implicit Encoding SLAM
by: Chen, Shi, et al.
Published: (2024)
by: Chen, Shi, et al.
Published: (2024)
When Sinks Help or Hurt: Unified Framework for Attention Sink in Large Vision-Language Models
by: Choi, Jiho, et al.
Published: (2026)
by: Choi, Jiho, et al.
Published: (2026)
TrafficBots V1.5: Traffic Simulation via Conditional VAEs and Transformers with Relative Pose Encoding
by: Zhang, Zhejun, et al.
Published: (2024)
by: Zhang, Zhejun, et al.
Published: (2024)
ReLoo: Reconstructing Humans Dressed in Loose Garments from Monocular Video in the Wild
by: Guo, Chen, et al.
Published: (2024)
by: Guo, Chen, et al.
Published: (2024)
MultiPly: Reconstruction of Multiple People from Monocular Video in the Wild
by: Jiang, Zeren, et al.
Published: (2024)
by: Jiang, Zeren, et al.
Published: (2024)
Exploration-Driven Generative Interactive Environments
by: Savov, Nedko, et al.
Published: (2025)
by: Savov, Nedko, et al.
Published: (2025)
From Scan to Action: Leveraging Realistic Scans for Embodied Scene Understanding
by: Halacheva, Anna-Maria, et al.
Published: (2025)
by: Halacheva, Anna-Maria, et al.
Published: (2025)
Condition-Invariant Semantic Segmentation
by: Sakaridis, Christos, et al.
Published: (2023)
by: Sakaridis, Christos, et al.
Published: (2023)
Enhanced Multi-Scale Cross-Attention for Person Image Generation
by: Tang, Hao, et al.
Published: (2025)
by: Tang, Hao, et al.
Published: (2025)
Towards Online Real-Time Memory-based Video Inpainting Transformers
by: Thiry, Guillaume, et al.
Published: (2024)
by: Thiry, Guillaume, et al.
Published: (2024)
Similar Items
-
Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation
by: Pasca, Razvan-George, et al.
Published: (2023) -
Leveraging Driver Field-of-View for Multimodal Ego-Trajectory Prediction
by: Akbiyik, M. Eren, et al.
Published: (2023) -
EgoGaussian: Dynamic Scene Understanding from Egocentric Video with 3D Gaussian Splatting
by: Zhang, Daiwei, et al.
Published: (2024) -
Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs
by: Kim, Sanghwan, et al.
Published: (2025) -
MOBIUS: Big-to-Mobile Universal Instance Segmentation via Multi-modal Bottleneck Fusion and Calibrated Decoder Pruning
by: Segu, Mattia, et al.
Published: (2025)