Aligning Actions and Walking to LLM-Generated Textual Descriptions
Fuente:
arXiv
Guardado en:
| Autores principales: | Chivereanu, Radu, Cosma, Adrian, Catruna, Andy, Rughinis, Razvan, Radoi, Emilian |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
The Paradox of Motion: Evidence for Spurious Correlations in Skeleton-based Gait Recognition Models
por: Cătrună, Andy, et al.
Publicado: (2024)
por: Cătrună, Andy, et al.
Publicado: (2024)
CrossGaze: A Strong Method for 3D Gaze Estimation in the Wild
por: Cătrună, Andy, et al.
Publicado: (2024)
por: Cătrună, Andy, et al.
Publicado: (2024)
MoME: Estimating Psychological Traits from Gait with Multi-Stage Mixture of Movement Experts
por: Cǎtrunǎ, Andy, et al.
Publicado: (2025)
por: Cǎtrunǎ, Andy, et al.
Publicado: (2025)
On Model and Data Scaling for Skeleton-based Self-Supervised Gait Recognition
por: Cosma, Adrian, et al.
Publicado: (2025)
por: Cosma, Adrian, et al.
Publicado: (2025)
GaitPT: Skeletons Are All You Need For Gait Recognition
por: Catruna, Andy, et al.
Publicado: (2023)
por: Catruna, Andy, et al.
Publicado: (2023)
Database-Agnostic Gait Enrollment using SetTransformers
por: Basoc, Nicoleta, et al.
Publicado: (2025)
por: Basoc, Nicoleta, et al.
Publicado: (2025)
Gait Recognition from Highly Compressed Videos
por: Niculae, Andrei, et al.
Publicado: (2024)
por: Niculae, Andrei, et al.
Publicado: (2024)
Spatial Colour Mixing Illusions as a Perception Stress Test for Vision-Language Models
por: Basoc, Nicoleta-Nina, et al.
Publicado: (2026)
por: Basoc, Nicoleta-Nina, et al.
Publicado: (2026)
What Makes a Good Doctor Response? A Study on Text-Based Telemedicine
por: Cosma, Adrian, et al.
Publicado: (2026)
por: Cosma, Adrian, et al.
Publicado: (2026)
A Retrieval-Based Approach to Medical Procedure Matching in Romanian
por: Niculae, Andrei, et al.
Publicado: (2025)
por: Niculae, Andrei, et al.
Publicado: (2025)
Vision-Free Retrieval: Rethinking Multimodal Search with Textual Scene Descriptions
por: Ntinou, Ioanna, et al.
Publicado: (2025)
por: Ntinou, Ioanna, et al.
Publicado: (2025)
MoReact: Generating Reactive Motion from Textual Descriptions
por: Xu, Xiyan, et al.
Publicado: (2025)
por: Xu, Xiyan, et al.
Publicado: (2025)
TextPSG: Panoptic Scene Graph Generation from Textual Descriptions
por: Zhao, Chengyang, et al.
Publicado: (2023)
por: Zhao, Chengyang, et al.
Publicado: (2023)
Learning to Generate Human-Human-Object Interactions from Textual Descriptions
por: Na, Jeonghyeon, et al.
Publicado: (2025)
por: Na, Jeonghyeon, et al.
Publicado: (2025)
RoMath: A Mathematical Reasoning Benchmark in Romanian
por: Cosma, Adrian, et al.
Publicado: (2024)
por: Cosma, Adrian, et al.
Publicado: (2024)
Contact-aware Human Motion Generation from Textual Descriptions
por: Ma, Sihan, et al.
Publicado: (2024)
por: Ma, Sihan, et al.
Publicado: (2024)
VisionTrap: Vision-Augmented Trajectory Prediction Guided by Textual Descriptions
por: Moon, Seokha, et al.
Publicado: (2024)
por: Moon, Seokha, et al.
Publicado: (2024)
Unbiasing through Textual Descriptions: Mitigating Representation Bias in Video Benchmarks
por: Shvetsova, Nina, et al.
Publicado: (2025)
por: Shvetsova, Nina, et al.
Publicado: (2025)
Motion Generation from Fine-grained Textual Descriptions
por: Li, Kunhang, et al.
Publicado: (2024)
por: Li, Kunhang, et al.
Publicado: (2024)
Zero-Shot Temporal Action Localization Through Textual Guidance
por: Liberatori, Benedetta, et al.
Publicado: (2026)
por: Liberatori, Benedetta, et al.
Publicado: (2026)
Training Language Models with homotokens Leads to Delayed Overfitting
por: Cosma, Adrian, et al.
Publicado: (2026)
por: Cosma, Adrian, et al.
Publicado: (2026)
Dr.Copilot: A Multi-Agent Prompt Optimized Assistant for Improving Patient-Doctor Communication in Romanian
por: Niculae, Andrei, et al.
Publicado: (2025)
por: Niculae, Andrei, et al.
Publicado: (2025)
The Strawberry Problem: Emergence of Character-level Understanding in Tokenized Language Models
por: Cosma, Adrian, et al.
Publicado: (2025)
por: Cosma, Adrian, et al.
Publicado: (2025)
Image Textualization: An Automatic Framework for Creating Accurate and Detailed Image Descriptions
por: Pi, Renjie, et al.
Publicado: (2024)
por: Pi, Renjie, et al.
Publicado: (2024)
Sat2Cap: Mapping Fine-Grained Textual Descriptions from Satellite Images
por: Dhakal, Aayush, et al.
Publicado: (2023)
por: Dhakal, Aayush, et al.
Publicado: (2023)
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
por: Sun, Zhengyang, et al.
Publicado: (2026)
por: Sun, Zhengyang, et al.
Publicado: (2026)
MMLNB: Multi-Modal Learning for Neuroblastoma Subtyping Classification Assisted with Textual Description Generation
por: Chen, Huangwei, et al.
Publicado: (2025)
por: Chen, Huangwei, et al.
Publicado: (2025)
Enhancing Action Recognition by Leveraging the Hierarchical Structure of Actions and Textual Context
por: Benavent-Lledo, Manuel, et al.
Publicado: (2024)
por: Benavent-Lledo, Manuel, et al.
Publicado: (2024)
Language-driven Description Generation and Common Sense Reasoning for Video Action Recognition
por: Hu, Xiaodan, et al.
Publicado: (2025)
por: Hu, Xiaodan, et al.
Publicado: (2025)
Distilling Textual Priors from LLM to Efficient Image Fusion
por: Zhang, Ran, et al.
Publicado: (2025)
por: Zhang, Ran, et al.
Publicado: (2025)
LAViTeR: Learning Aligned Visual and Textual Representations Assisted by Image and Caption Generation
por: Hashemi, Mohammad Abuzar, et al.
Publicado: (2021)
por: Hashemi, Mohammad Abuzar, et al.
Publicado: (2021)
ActionHub: A Large-scale Action Video Description Dataset for Zero-shot Action Recognition
por: Zhou, Jiaming, et al.
Publicado: (2024)
por: Zhou, Jiaming, et al.
Publicado: (2024)
Vision-and-Language Navigation with Analogical Textual Descriptions in LLMs
por: Zhang, Yue, et al.
Publicado: (2025)
por: Zhang, Yue, et al.
Publicado: (2025)
Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision
por: Yoshida, Tomoya, et al.
Publicado: (2025)
por: Yoshida, Tomoya, et al.
Publicado: (2025)
Instance-aware Image Colorization with Controllable Textual Descriptions and Segmentation Masks
por: An, Yanru, et al.
Publicado: (2025)
por: An, Yanru, et al.
Publicado: (2025)
Teaching Time Series to See and Speak: Forecasting with Aligned Visual and Textual Perspectives
por: Dong, Sixun, et al.
Publicado: (2025)
por: Dong, Sixun, et al.
Publicado: (2025)
OrienText: Surface Oriented Textual Image Generation
por: Paliwal, Shubham Singh, et al.
Publicado: (2025)
por: Paliwal, Shubham Singh, et al.
Publicado: (2025)
Precise Parameter Localization for Textual Generation in Diffusion Models
por: Staniszewski, Łukasz, et al.
Publicado: (2025)
por: Staniszewski, Łukasz, et al.
Publicado: (2025)
AlignedGen: Aligning Style Across Generated Images
por: Zhang, Jiexuan, et al.
Publicado: (2025)
por: Zhang, Jiexuan, et al.
Publicado: (2025)
MAGR: Manifold-Aligned Graph Regularization for Continual Action Quality Assessment
por: Zhou, Kanglei, et al.
Publicado: (2024)
por: Zhou, Kanglei, et al.
Publicado: (2024)
Ejemplares similares
-
The Paradox of Motion: Evidence for Spurious Correlations in Skeleton-based Gait Recognition Models
por: Cătrună, Andy, et al.
Publicado: (2024) -
CrossGaze: A Strong Method for 3D Gaze Estimation in the Wild
por: Cătrună, Andy, et al.
Publicado: (2024) -
MoME: Estimating Psychological Traits from Gait with Multi-Stage Mixture of Movement Experts
por: Cǎtrunǎ, Andy, et al.
Publicado: (2025) -
On Model and Data Scaling for Skeleton-based Self-Supervised Gait Recognition
por: Cosma, Adrian, et al.
Publicado: (2025) -
GaitPT: Skeletons Are All You Need For Gait Recognition
por: Catruna, Andy, et al.
Publicado: (2023)