Gespeichert in:
| Hauptverfasser: | Lai, Bolin, Dai, Xiaoliang, Chen, Lawrence, Pang, Guan, Rehg, James M., Liu, Miao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2312.03849 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
von: Lai, Bolin, et al.
Veröffentlicht: (2023)
von: Lai, Bolin, et al.
Veröffentlicht: (2023)
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
von: Lai, Bolin, et al.
Veröffentlicht: (2022)
von: Lai, Bolin, et al.
Veröffentlicht: (2022)
Learning Predictive Visuomotor Coordination
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)
von: Jia, Wenqi, et al.
Veröffentlicht: (2025)
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation
von: Lai, Bolin, et al.
Veröffentlicht: (2024)
von: Lai, Bolin, et al.
Veröffentlicht: (2024)
Human Action Anticipation: A Survey
von: Lai, Bolin, et al.
Veröffentlicht: (2024)
von: Lai, Bolin, et al.
Veröffentlicht: (2024)
Towards Online Multi-Modal Social Interaction Understanding
von: Li, Xinpeng, et al.
Veröffentlicht: (2025)
von: Li, Xinpeng, et al.
Veröffentlicht: (2025)
AMEGO: Active Memory from long EGOcentric videos
von: Goletto, Gabriele, et al.
Veröffentlicht: (2024)
von: Goletto, Gabriele, et al.
Veröffentlicht: (2024)
PREGO: online mistake detection in PRocedural EGOcentric videos
von: Flaborea, Alessandro, et al.
Veröffentlicht: (2024)
von: Flaborea, Alessandro, et al.
Veröffentlicht: (2024)
Modeling Multimodal Social Interactions: New Challenges and Baselines with Densely Aligned Representations
von: Lee, Sangmin, et al.
Veröffentlicht: (2024)
von: Lee, Sangmin, et al.
Veröffentlicht: (2024)
SocialGesture: Delving into Multi-person Gesture Understanding
von: Cao, Xu, et al.
Veröffentlicht: (2025)
von: Cao, Xu, et al.
Veröffentlicht: (2025)
Toward Diffusible High-Dimensional Latent Spaces: A Frequency Perspective
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
von: Lai, Bolin, et al.
Veröffentlicht: (2025)
Omni-MMSI: Toward Identity-attributed Social Interaction Understanding
von: Li, Xinpeng, et al.
Veröffentlicht: (2026)
von: Li, Xinpeng, et al.
Veröffentlicht: (2026)
Generative Visual Instruction Tuning
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2024)
von: Hernandez, Jefferson, et al.
Veröffentlicht: (2024)
TI-PREGO: Chain of Thought and In-Context Learning for Online Mistake Detection in PRocedural EGOcentric Videos
von: Plini, Leonardo, et al.
Veröffentlicht: (2024)
von: Plini, Leonardo, et al.
Veröffentlicht: (2024)
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions
von: Kim, Junho, et al.
Veröffentlicht: (2026)
von: Kim, Junho, et al.
Veröffentlicht: (2026)
What is the Visual Cognition Gap between Humans and Multimodal LLMs?
von: Cao, Xu, et al.
Veröffentlicht: (2024)
von: Cao, Xu, et al.
Veröffentlicht: (2024)
InstructVLA: Vision-Language-Action Instruction Tuning from Understanding to Manipulation
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
von: Yang, Shuai, et al.
Veröffentlicht: (2025)
The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective
von: Jia, Wenqi, et al.
Veröffentlicht: (2023)
von: Jia, Wenqi, et al.
Veröffentlicht: (2023)
DiffEye: Diffusion-Based Continuous Eye-Tracking Data Generation Conditioned on Natural Images
von: Kara, Ozgur, et al.
Veröffentlicht: (2025)
von: Kara, Ozgur, et al.
Veröffentlicht: (2025)
Personalized Visual Instruction Tuning
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
von: Pi, Renjie, et al.
Veröffentlicht: (2024)
Comparison Visual Instruction Tuning
von: Lin, Wei, et al.
Veröffentlicht: (2024)
von: Lin, Wei, et al.
Veröffentlicht: (2024)
Visual Instruction Tuning with Chain of Region-of-Interest
von: Chen, Yixin, et al.
Veröffentlicht: (2025)
von: Chen, Yixin, et al.
Veröffentlicht: (2025)
MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs
von: Ye, Wenqian, et al.
Veröffentlicht: (2024)
von: Ye, Wenqian, et al.
Veröffentlicht: (2024)
LEGO-Motion: Learning-Enhanced Grids with Occupancy Instance Modeling for Class-Agnostic Motion Prediction
von: Qian, Kangan, et al.
Veröffentlicht: (2025)
von: Qian, Kangan, et al.
Veröffentlicht: (2025)
Immune2V: Image Immunization Against Dual-Stream Image-to-Video Generation
von: Long, Zeqian, et al.
Veröffentlicht: (2026)
von: Long, Zeqian, et al.
Veröffentlicht: (2026)
Symmetry Strikes Back: From Single-Image Symmetry Detection to 3D Generation
von: Li, Xiang, et al.
Veröffentlicht: (2024)
von: Li, Xiang, et al.
Veröffentlicht: (2024)
INST-IT: Boosting Instance Understanding via Explicit Visual Prompt Instruction Tuning
von: Peng, Wujian, et al.
Veröffentlicht: (2024)
von: Peng, Wujian, et al.
Veröffentlicht: (2024)
Cue3D: Quantifying the Role of Image Cues in Single-Image 3D Generation
von: Li, Xiang, et al.
Veröffentlicht: (2025)
von: Li, Xiang, et al.
Veröffentlicht: (2025)
MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
von: Tong, Shengbang, et al.
Veröffentlicht: (2024)
Instructify: Demystifying Metadata to Visual Instruction Tuning Data Conversion
von: Hansen, Jacob, et al.
Veröffentlicht: (2025)
von: Hansen, Jacob, et al.
Veröffentlicht: (2025)
Gaze-LLE: Gaze Target Estimation via Large-Scale Learned Encoders
von: Ryan, Fiona, et al.
Veröffentlicht: (2024)
von: Ryan, Fiona, et al.
Veröffentlicht: (2024)
LEGO: Self-Supervised Representation Learning for Scene Text Images
von: Ren, Yujin, et al.
Veröffentlicht: (2024)
von: Ren, Yujin, et al.
Veröffentlicht: (2024)
Noise-Tolerant Learning for Audio-Visual Action Recognition
von: Han, Haochen, et al.
Veröffentlicht: (2022)
von: Han, Haochen, et al.
Veröffentlicht: (2022)
Osprey: Pixel Understanding with Visual Instruction Tuning
von: Yuan, Yuqian, et al.
Veröffentlicht: (2023)
von: Yuan, Yuqian, et al.
Veröffentlicht: (2023)
Pose-RFT: Enhancing MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning
von: Li, Bao, et al.
Veröffentlicht: (2025)
von: Li, Bao, et al.
Veröffentlicht: (2025)
LLARVA: Vision-Action Instruction Tuning Enhances Robot Learning
von: Niu, Dantong, et al.
Veröffentlicht: (2024)
von: Niu, Dantong, et al.
Veröffentlicht: (2024)
Instruction Tuning-free Visual Token Complement for Multimodal LLMs
von: Wang, Dongsheng, et al.
Veröffentlicht: (2024)
von: Wang, Dongsheng, et al.
Veröffentlicht: (2024)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
von: Li, Hao, et al.
Veröffentlicht: (2025)
von: Li, Hao, et al.
Veröffentlicht: (2025)
Medical Image Understanding Improves Survival Prediction via Visual Instruction Tuning
von: Liu, Xixi, et al.
Veröffentlicht: (2026)
von: Liu, Xixi, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation
von: Lai, Bolin, et al.
Veröffentlicht: (2023) -
In the Eye of Transformer: Global-Local Correlation for Egocentric Gaze Estimation
von: Lai, Bolin, et al.
Veröffentlicht: (2022) -
Learning Predictive Visuomotor Coordination
von: Jia, Wenqi, et al.
Veröffentlicht: (2025) -
Unified Text-Image-to-Video Generation: A Training-Free Approach to Flexible Visual Conditioning
von: Lai, Bolin, et al.
Veröffentlicht: (2025) -
Unleashing In-context Learning of Autoregressive Models for Few-shot Image Manipulation
von: Lai, Bolin, et al.
Veröffentlicht: (2024)