Pose Priors from Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Subramanian, Sanjay, Ng, Evonne, Müller, Lea, Klein, Dan, Ginosar, Shiry, Darrell, Trevor |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Recursive Visual Programming
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023)
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
von: Shang, Chuyi, et al.
Veröffentlicht: (2024)
von: Shang, Chuyi, et al.
Veröffentlicht: (2024)
Synergy and Synchrony in Couple Dances
von: Maluleke, Vongani, et al.
Veröffentlicht: (2024)
von: Maluleke, Vongani, et al.
Veröffentlicht: (2024)
Poly-Autoregressive Prediction for Modeling Interactions
von: Thakkar, Neerja, et al.
Veröffentlicht: (2025)
von: Thakkar, Neerja, et al.
Veröffentlicht: (2025)
Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
Vision-Language Models Create Cross-Modal Task Representations
von: Luo, Grace, et al.
Veröffentlicht: (2024)
von: Luo, Grace, et al.
Veröffentlicht: (2024)
KiVA: Kid-inspired Visual Analogies for Testing Large Multimodal Models
von: Yiu, Eunice, et al.
Veröffentlicht: (2024)
von: Yiu, Eunice, et al.
Veröffentlicht: (2024)
AutoPresent: Designing Structured Visuals from Scratch
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
von: Ge, Jiaxin, et al.
Veröffentlicht: (2025)
From Audio to Photoreal Embodiment: Synthesizing Humans in Conversations
von: Ng, Evonne, et al.
Veröffentlicht: (2024)
von: Ng, Evonne, et al.
Veröffentlicht: (2024)
Diffusion Models as Data Mining Tools
von: Siglidis, Ioannis, et al.
Veröffentlicht: (2024)
von: Siglidis, Ioannis, et al.
Veröffentlicht: (2024)
Forecasting Motion in the Wild
von: Thakkar, Neerja, et al.
Veröffentlicht: (2026)
von: Thakkar, Neerja, et al.
Veröffentlicht: (2026)
LLM-grounded Video Diffusion Models
von: Lian, Long, et al.
Veröffentlicht: (2023)
von: Lian, Long, et al.
Veröffentlicht: (2023)
Back into Plato's Cave: Examining Cross-modal Representational Convergence at Scale
von: Koepke, A. Sophia, et al.
Veröffentlicht: (2026)
von: Koepke, A. Sophia, et al.
Veröffentlicht: (2026)
Gaussian Masked Autoencoders
von: Rajasegaran, Jathushan, et al.
Veröffentlicht: (2025)
von: Rajasegaran, Jathushan, et al.
Veröffentlicht: (2025)
Dual-Process Image Generation
von: Luo, Grace, et al.
Veröffentlicht: (2025)
von: Luo, Grace, et al.
Veröffentlicht: (2025)
Mechanistic Finetuning of Vision-Language-Action Models via Few-Shot Demonstrations
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2025)
Ham2Pose: Animating Sign Language Notation into Pose Sequences
von: Shalev-Arkushin, Rotem, et al.
Veröffentlicht: (2022)
von: Shalev-Arkushin, Rotem, et al.
Veröffentlicht: (2022)
Compositional Chain-of-Thought Prompting for Large Multimodal Models
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
von: Girrbach, Leander, et al.
Veröffentlicht: (2025)
von: Girrbach, Leander, et al.
Veröffentlicht: (2025)
Pose-Based Sign Language Appearance Transfer
von: Moryossef, Amit, et al.
Veröffentlicht: (2024)
von: Moryossef, Amit, et al.
Veröffentlicht: (2024)
Enhancing Few-Shot Vision-Language Classification with Large Multimodal Model Features
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
von: Mitra, Chancharik, et al.
Veröffentlicht: (2024)
Diffusion Forcing for Multi-Agent Interaction Sequence Modeling
von: Maluleke, Vongani H., et al.
Veröffentlicht: (2025)
von: Maluleke, Vongani H., et al.
Veröffentlicht: (2025)
Puzzled by Puzzles: When Vision-Language Models Can't Take a Hint
von: Lee, Heekyung, et al.
Veröffentlicht: (2025)
von: Lee, Heekyung, et al.
Veröffentlicht: (2025)
Expanding the Boundaries of Vision Prior Knowledge in Multi-modal Large Language Models
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
von: Liang, Qiao, et al.
Veröffentlicht: (2025)
LLM-grounded Diffusion: Enhancing Prompt Understanding of Text-to-Image Diffusion Models with Large Language Models
von: Lian, Long, et al.
Veröffentlicht: (2023)
von: Lian, Long, et al.
Veröffentlicht: (2023)
Revisiting the Role of Language Priors in Vision-Language Models
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2023)
von: Lin, Zhiqiu, et al.
Veröffentlicht: (2023)
Analyzing The Language of Visual Tokens
von: Chan, David M., et al.
Veröffentlicht: (2024)
von: Chan, David M., et al.
Veröffentlicht: (2024)
Seeing Isn't Believing: Uncovering Blind Spots in Evaluator Vision-Language Models
von: Khan, Mohammed Safi Ur Rahman, et al.
Veröffentlicht: (2026)
von: Khan, Mohammed Safi Ur Rahman, et al.
Veröffentlicht: (2026)
Pose-Based Sign Language Spotting via an End-to-End Encoder Architecture
von: Johnny, Samuel Ebimobowei, et al.
Veröffentlicht: (2025)
von: Johnny, Samuel Ebimobowei, et al.
Veröffentlicht: (2025)
A Review on Large Language Models for Visual Analytics
von: Agarwal, Navya Sonal, et al.
Veröffentlicht: (2025)
von: Agarwal, Navya Sonal, et al.
Veröffentlicht: (2025)
Describing Differences in Image Sets with Natural Language
von: Dunlap, Lisa, et al.
Veröffentlicht: (2023)
von: Dunlap, Lisa, et al.
Veröffentlicht: (2023)
Overcoming Language Priors for Visual Question Answering Based on Knowledge Distillation
von: Peng, Daowan, et al.
Veröffentlicht: (2025)
von: Peng, Daowan, et al.
Veröffentlicht: (2025)
VLind-Bench: Measuring Language Priors in Large Vision-Language Models
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
von: Lee, Kang-il, et al.
Veröffentlicht: (2024)
Enabling Stroke-Level Structural Analysis of Hieroglyphic Scripts without Language-Specific Priors
von: Luo, Fuwen, et al.
Veröffentlicht: (2026)
von: Luo, Fuwen, et al.
Veröffentlicht: (2026)
Lost in Space? Vision-Language Models Struggle with Relative Camera Pose Estimation
von: Deng, Ken, et al.
Veröffentlicht: (2026)
von: Deng, Ken, et al.
Veröffentlicht: (2026)
Readout Guidance: Learning Control from Diffusion Features
von: Luo, Grace, et al.
Veröffentlicht: (2023)
von: Luo, Grace, et al.
Veröffentlicht: (2023)
DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images
von: Baral, Sami, et al.
Veröffentlicht: (2025)
von: Baral, Sami, et al.
Veröffentlicht: (2025)
READ: Recurrent Adapter with Partial Video-Language Alignment for Parameter-Efficient Transfer Learning in Low-Resource Video-Language Modeling
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
DemaFormer: Damped Exponential Moving Average Transformer with Energy-Based Modeling for Temporal Language Grounding
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
von: Nguyen, Thong, et al.
Veröffentlicht: (2023)
From Sight to Insight: Improving Visual Reasoning Capabilities of Multimodal Models via Reinforcement Learning
von: Sharif, Omar, et al.
Veröffentlicht: (2026)
von: Sharif, Omar, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Recursive Visual Programming
von: Ge, Jiaxin, et al.
Veröffentlicht: (2023) -
TraveLER: A Modular Multi-LMM Agent Framework for Video Question-Answering
von: Shang, Chuyi, et al.
Veröffentlicht: (2024) -
Synergy and Synchrony in Couple Dances
von: Maluleke, Vongani, et al.
Veröffentlicht: (2024) -
Poly-Autoregressive Prediction for Modeling Interactions
von: Thakkar, Neerja, et al.
Veröffentlicht: (2025) -
Which One? Leveraging Context Between Objects and Multiple Views for Language Grounding
von: Mitra, Chancharik, et al.
Veröffentlicht: (2023)