Learning Multiple Object States from Actions via Large Language Models
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tateno, Masatoshi, Yagi, Takuma, Furuta, Ryosuke, Sato, Yoichi |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
von: Tateno, Masatoshi, et al.
Veröffentlicht: (2025)
von: Tateno, Masatoshi, et al.
Veröffentlicht: (2025)
Seeking Flat Minima with Mean Teacher on Semi- and Weakly-Supervised Domain Generalization for Object Detection
von: Furuta, Ryosuke, et al.
Veröffentlicht: (2023)
von: Furuta, Ryosuke, et al.
Veröffentlicht: (2023)
ActionVOS: Actions as Prompts for Video Object Segmentation
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2024)
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2024)
FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation
von: Yagi, Takuma, et al.
Veröffentlicht: (2024)
von: Yagi, Takuma, et al.
Veröffentlicht: (2024)
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
von: Ohkawa, Takehiko, et al.
Veröffentlicht: (2023)
von: Ohkawa, Takehiko, et al.
Veröffentlicht: (2023)
EgoInstruct: An Egocentric Video Dataset of Face-to-face Instructional Interactions with Multi-modal LLM Benchmarking
von: Sakai, Yuki, et al.
Veröffentlicht: (2025)
von: Sakai, Yuki, et al.
Veröffentlicht: (2025)
Pre-Training for 3D Hand Pose Estimation with Contrastive Learning on Large-Scale Hand Images in the Wild
von: Lin, Nie, et al.
Veröffentlicht: (2024)
von: Lin, Nie, et al.
Veröffentlicht: (2024)
AssemblyHands-X: Modeling 3D Hand-Body Coordination for Understanding Bimanual Human Activities
von: Banno, Tatsuro, et al.
Veröffentlicht: (2025)
von: Banno, Tatsuro, et al.
Veröffentlicht: (2025)
Learning Gaussian Data Augmentation in Feature Space for One-shot Object Detection in Manga
von: Taniguchi, Takara, et al.
Veröffentlicht: (2024)
von: Taniguchi, Takara, et al.
Veröffentlicht: (2024)
Inference-time Trajectory Optimization for Manga Image Editing
von: Furuta, Ryosuke
Veröffentlicht: (2026)
von: Furuta, Ryosuke
Veröffentlicht: (2026)
Leadership Assessment in Pediatric Intensive Care Unit Team Training
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
Affordance-Guided Diffusion Prior for 3D Hand Reconstruction
von: Suzuki, Naru, et al.
Veröffentlicht: (2025)
von: Suzuki, Naru, et al.
Veröffentlicht: (2025)
Multi-speaker Attention Alignment for Multimodal Social Interaction
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2025)
SiMHand: Mining Similar Hands for Large-Scale 3D Hand Pose Pre-training
von: Lin, Nie, et al.
Veröffentlicht: (2025)
von: Lin, Nie, et al.
Veröffentlicht: (2025)
EgoBrain: Synergizing Minds and Eyes For Human Action Understanding
von: Lin, Nie, et al.
Veröffentlicht: (2025)
von: Lin, Nie, et al.
Veröffentlicht: (2025)
Bidirectional Action Sequence Learning for Long-term Action Anticipation with Large Language Models
von: Sato, Yuji, et al.
Veröffentlicht: (2025)
von: Sato, Yuji, et al.
Veröffentlicht: (2025)
Generative Modeling of Shape-Dependent Self-Contact Human Poses
von: Ohkawa, Takehiko, et al.
Veröffentlicht: (2025)
von: Ohkawa, Takehiko, et al.
Veröffentlicht: (2025)
Egocentric Gaze Estimation via Neck-Mounted Camera
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
von: Huang, Haoyu, et al.
Veröffentlicht: (2026)
Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance
von: Zhang, Mingfang, et al.
Veröffentlicht: (2025)
von: Zhang, Mingfang, et al.
Veröffentlicht: (2025)
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition
von: Zhang, Mingfang, et al.
Veröffentlicht: (2024)
von: Zhang, Mingfang, et al.
Veröffentlicht: (2024)
Generative Hierarchical Temporal Transformer for Hand Pose and Action Modeling
von: Wen, Yilin, et al.
Veröffentlicht: (2023)
von: Wen, Yilin, et al.
Veröffentlicht: (2023)
Extending Dataset Pruning to Object Detection: A Variance-based Approach
von: Yagi, Ryota
Veröffentlicht: (2025)
von: Yagi, Ryota
Veröffentlicht: (2025)
VSRD++: Autolabeling for 3D Object Detection via Instance-Aware Volumetric Silhouette Rendering
von: Liu, Zihua, et al.
Veröffentlicht: (2025)
von: Liu, Zihua, et al.
Veröffentlicht: (2025)
VISA: Reasoning Video Object Segmentation via Large Language Models
von: Yan, Cilin, et al.
Veröffentlicht: (2024)
von: Yan, Cilin, et al.
Veröffentlicht: (2024)
Zero-shot Action Localization via the Confidence of Large Vision-Language Models
von: Aklilu, Josiah, et al.
Veröffentlicht: (2024)
von: Aklilu, Josiah, et al.
Veröffentlicht: (2024)
The N-Body Problem: Parallel Execution from Single-Person Egocentric Video
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
von: Zhu, Zhifan, et al.
Veröffentlicht: (2025)
Learning to Reason and Navigate: Parameter Efficient Action Planning with Large Language Models
von: Mohammadi, Bahram, et al.
Veröffentlicht: (2025)
von: Mohammadi, Bahram, et al.
Veröffentlicht: (2025)
SFHand: Learning Embodied Manipulation by Streaming Egocentric 3D Hand Forecasting
von: Liu, Ruicong, et al.
Veröffentlicht: (2025)
von: Liu, Ruicong, et al.
Veröffentlicht: (2025)
CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model
von: Yao, Ziyu, et al.
Veröffentlicht: (2025)
von: Yao, Ziyu, et al.
Veröffentlicht: (2025)
Weakly Supervised Temporal Action Localization via Dual-Prior Collaborative Learning Guided by Multimodal Large Language Models
von: Zhang, Quan, et al.
Veröffentlicht: (2024)
von: Zhang, Quan, et al.
Veröffentlicht: (2024)
Language-guided Learning for Object Detection Tackling Multiple Variations in Aerial Images
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
von: Park, Sungjune, et al.
Veröffentlicht: (2025)
VSRD: Instance-Aware Volumetric Silhouette Rendering for Weakly Supervised 3D Object Detection
von: Liu, Zihua, et al.
Veröffentlicht: (2024)
von: Liu, Zihua, et al.
Veröffentlicht: (2024)
MambaTrack: A Simple Baseline for Multiple Object Tracking with State Space Model
von: Xiao, Changcheng, et al.
Veröffentlicht: (2024)
von: Xiao, Changcheng, et al.
Veröffentlicht: (2024)
FairT2I: Mitigating Social Bias in Text-to-Image Generation via Large Language Model-Assisted Detection and Attribute Rebalancing
von: Sakurai, Jinya, et al.
Veröffentlicht: (2025)
von: Sakurai, Jinya, et al.
Veröffentlicht: (2025)
Single-to-Dual-View Adaptation for Egocentric 3D Hand Pose Estimation
von: Liu, Ruicong, et al.
Veröffentlicht: (2024)
von: Liu, Ruicong, et al.
Veröffentlicht: (2024)
Motion State: A New Benchmark Multiple Object Tracking
von: Feng, Yang, et al.
Veröffentlicht: (2023)
von: Feng, Yang, et al.
Veröffentlicht: (2023)
LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model
von: Gao, Quankai, et al.
Veröffentlicht: (2026)
von: Gao, Quankai, et al.
Veröffentlicht: (2026)
Point Linguist Model: Segment Any Object via Bridged Large 3D-Language Model
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2025)
von: Huang, Zhuoxu, et al.
Veröffentlicht: (2025)
Nullu: Mitigating Object Hallucinations in Large Vision-Language Models via HalluSpace Projection
von: Yang, Le, et al.
Veröffentlicht: (2024)
von: Yang, Le, et al.
Veröffentlicht: (2024)
YOLIC: An Efficient Method for Object Localization and Classification on Edge Devices
von: Su, Kai, et al.
Veröffentlicht: (2023)
von: Su, Kai, et al.
Veröffentlicht: (2023)
Ähnliche Einträge
-
HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
von: Tateno, Masatoshi, et al.
Veröffentlicht: (2025) -
Seeking Flat Minima with Mean Teacher on Semi- and Weakly-Supervised Domain Generalization for Object Detection
von: Furuta, Ryosuke, et al.
Veröffentlicht: (2023) -
ActionVOS: Actions as Prompts for Video Object Segmentation
von: Ouyang, Liangyang, et al.
Veröffentlicht: (2024) -
FineBio: A Fine-Grained Video Dataset of Biological Experiments with Hierarchical Annotation
von: Yagi, Takuma, et al.
Veröffentlicht: (2024) -
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos
von: Ohkawa, Takehiko, et al.
Veröffentlicht: (2023)