Saved in:
| Main Authors: | Lai, Bolin, Juefei-Xu, Felix, Liu, Miao, Dai, Xiaoliang, Mehta, Nikhil, Zhu, Chenguang, Huang, Zeyi, Rehg, James M., Lee, Sangmin, Zhang, Ning, Xiao, Tong |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2412.01027 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
by: Lai, Bolin, et al.
Published: (2025)
by: Lai, Bolin, et al.
Published: (2025)
LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning
by: Lai, Bolin, et al.
Published: (2023)
by: Lai, Bolin, et al.
Published: (2023)
Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
by: Lee, Sangmin, et al.
Published: (2024)
by: Lee, Sangmin, et al.
Published: (2024)
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
by: Lai, Bolin, et al.
Published: (2022)
by: Lai, Bolin, et al.
Published: (2022)
SocialGesture: Delving into Multi-person Gesture Understanding
by: Cao, Xu, et al.
Published: (2025)
by: Cao, Xu, et al.
Published: (2025)
Learning Predictive Visuomotor Coordination
by: Jia, Wenqi, et al.
Published: (2025)
by: Jia, Wenqi, et al.
Published: (2025)
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
by: Lai, Bolin, et al.
Published: (2023)
by: Lai, Bolin, et al.
Published: (2023)
Building a Mind Palace: Structuring Environment-Grounded Semantic Graphs for Effective Long Video Analysis with LLMs
by: Huang, Zeyi, et al.
Published: (2025)
by: Huang, Zeyi, et al.
Published: (2025)
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions
by: Kim, Junho, et al.
Published: (2026)
by: Kim, Junho, et al.
Published: (2026)
Improving Chain-of-Thought Efficiency for Autoregressive Image Generation
by: Gu, Zeqi, et al.
Published: (2025)
by: Gu, Zeqi, et al.
Published: (2025)
Towards Social AI: A Survey on Understanding Social Interactions
by: Lee, Sangmin, et al.
Published: (2024)
by: Lee, Sangmin, et al.
Published: (2024)
Towards Online Multi-Modal Social Interaction Understanding
by: Li, Xinpeng, et al.
Published: (2025)
by: Li, Xinpeng, et al.
Published: (2025)
Unleashing the Potential of the Diffusion Model in Few-shot Semantic Segmentation
by: Zhu, Muzhi, et al.
Published: (2024)
by: Zhu, Muzhi, et al.
Published: (2024)
Gaze-LLE: Gaze Target Estimation via Large-Scale Learned Encoders
by: Ryan, Fiona, et al.
Published: (2024)
by: Ryan, Fiona, et al.
Published: (2024)
Differentially Private In-context Learning via Sampling Few-shot Mixed with Zero-shot Outputs
by: Flemings, James, et al.
Published: (2025)
by: Flemings, James, et al.
Published: (2025)
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
by: Lai, Bolin, et al.
Published: (2025)
by: Lai, Bolin, et al.
Published: (2025)
Apollo: An Exploration of Video Understanding in Large Multimodal Models
by: Zohar, Orr, et al.
Published: (2024)
by: Zohar, Orr, et al.
Published: (2024)
Few-shot Protein Fitness Prediction via In-context Learning and Test-time Training
by: Teufel, Felix, et al.
Published: (2025)
by: Teufel, Felix, et al.
Published: (2025)
Leveraging Object Priors for Point Tracking
by: Boote, Bikram, et al.
Published: (2024)
by: Boote, Bikram, et al.
Published: (2024)
Unleashing the Power of Meta-tuning for Few-shot Generalization Through Sparse Interpolated Experts
by: Chen, Shengzhuang, et al.
Published: (2024)
by: Chen, Shengzhuang, et al.
Published: (2024)
ICPL: Few-shot In-context Preference Learning via LLMs
by: Yu, Chao, et al.
Published: (2024)
by: Yu, Chao, et al.
Published: (2024)
Human Action Anticipation: A Survey
by: Lai, Bolin, et al.
Published: (2024)
by: Lai, Bolin, et al.
Published: (2024)
ZeroShape: Regression-based Zero-shot Shape Reconstruction
by: Huang, Zixuan, et al.
Published: (2023)
by: Huang, Zixuan, et al.
Published: (2023)
Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
by: Li, Xinpeng, et al.
Published: (2026)
by: Li, Xinpeng, et al.
Published: (2026)
FAR-Dex: Few-shot Data Augmentation and Adaptive Residual Policy Refinement for Dexterous Manipulation
by: Bai, Yushan, et al.
Published: (2026)
by: Bai, Yushan, et al.
Published: (2026)
Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs
by: Zhang, Haochen, et al.
Published: (2026)
by: Zhang, Haochen, et al.
Published: (2026)
Intent-driven In-context Learning for Few-shot Dialogue State Tracking
by: Yi, Zihao, et al.
Published: (2024)
by: Yi, Zihao, et al.
Published: (2024)
Token-Shuffle: Towards High-Resolution Image Generation with Autoregressive Models
by: Ma, Xu, et al.
Published: (2025)
by: Ma, Xu, et al.
Published: (2025)
One-shot In-context Part Segmentation
by: Dai, Zhenqi, et al.
Published: (2025)
by: Dai, Zhenqi, et al.
Published: (2025)
MediCLIP: Adapting CLIP for Few-shot Medical Image Anomaly Detection
by: Zhang, Ximiao, et al.
Published: (2024)
by: Zhang, Ximiao, et al.
Published: (2024)
Make-it-Real: Unleashing Large Multimodal Model for Painting 3D Objects with Realistic Materials
by: Fang, Ye, et al.
Published: (2024)
by: Fang, Ye, et al.
Published: (2024)
Generating Synthetic Datasets for Few-shot Prompt Tuning
by: Guo, Xu, et al.
Published: (2024)
by: Guo, Xu, et al.
Published: (2024)
Cross-domain Few-shot In-context Learning for Enhancing Traffic Sign Recognition
by: Gan, Yaozong, et al.
Published: (2024)
by: Gan, Yaozong, et al.
Published: (2024)
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
by: Ye, Wenqian, et al.
Published: (2024)
by: Ye, Wenqian, et al.
Published: (2024)
DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural Images
by: Kara, Ozgur, et al.
Published: (2025)
by: Kara, Ozgur, et al.
Published: (2025)
DynaIP: Dynamic Image Prompt Adapter for Scalable Zero-shot Personalized Text-to-Image Generation
by: Wang, Zhizhong, et al.
Published: (2025)
by: Wang, Zhizhong, et al.
Published: (2025)
Towards Few-shot Entity Recognition in Document Images: A Graph Neural Network Approach Robust to Image Manipulation
by: Krishnan, Prashant, et al.
Published: (2023)
by: Krishnan, Prashant, et al.
Published: (2023)
Retrieval-augmented Few-shot Medical Image Segmentation with Foundation Models
by: Zhao, Lin, et al.
Published: (2024)
by: Zhao, Lin, et al.
Published: (2024)
Unleashing Large Language Models' Proficiency in Zero-shot Essay Scoring
by: Lee, Sanwoo, et al.
Published: (2024)
by: Lee, Sanwoo, et al.
Published: (2024)
Cue3D: Quantifying the Role of Image Cues in Single-Image 3D Generation
by: Li, Xiang, et al.
Published: (2025)
by: Li, Xiang, et al.
Published: (2025)
Similar Items
-
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
by: Lai, Bolin, et al.
Published: (2025) -
LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning
by: Lai, Bolin, et al.
Published: (2023) -
Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
by: Lee, Sangmin, et al.
Published: (2024) -
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
by: Lai, Bolin, et al.
Published: (2022) -
SocialGesture: Delving into Multi-person Gesture Understanding
by: Cao, Xu, et al.
Published: (2025)