Language-Informed Visual Concept Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Lee, Sharon, Zhang, Yunzhi, Wu, Shangzhe, Wu, Jiajun |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Seeing a Rose in Five Thousand Ways
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2022)
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2022)
The Scene Language: Representing Scenes with Programs, Words, and Embeddings
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2024)
Ponymation: Learning Articulated 3D Animal Motions from Unlabeled Online Videos
von: Sun, Keqiang, et al.
Veröffentlicht: (2023)
von: Sun, Keqiang, et al.
Veröffentlicht: (2023)
Birth and Death of a Rose
von: Geng, Chen, et al.
Veröffentlicht: (2024)
von: Geng, Chen, et al.
Veröffentlicht: (2024)
NeuROK: Generative 4D Neural Object Kinematics
von: Geng, Chen, et al.
Veröffentlicht: (2026)
von: Geng, Chen, et al.
Veröffentlicht: (2026)
Stanford-ORB: A Real-World 3D Object Inverse Rendering Benchmark
von: Kuang, Zhengfei, et al.
Veröffentlicht: (2023)
von: Kuang, Zhengfei, et al.
Veröffentlicht: (2023)
Learning the 3D Fauna of the Web
von: Li, Zizhang, et al.
Veröffentlicht: (2024)
von: Li, Zizhang, et al.
Veröffentlicht: (2024)
Choreographing a World of Dynamic Objects
von: Lyu, Yanzhe, et al.
Veröffentlicht: (2026)
von: Lyu, Yanzhe, et al.
Veröffentlicht: (2026)
Anymate: A Dataset and Baselines for Learning 3D Object Rigging
von: Deng, Yufan, et al.
Veröffentlicht: (2025)
von: Deng, Yufan, et al.
Veröffentlicht: (2025)
Revisiting Multi-Task Visual Representation Learning
von: Di, Shangzhe, et al.
Veröffentlicht: (2026)
von: Di, Shangzhe, et al.
Veröffentlicht: (2026)
Category-Agnostic Neural Object Rigging
von: He, Guangzhao, et al.
Veröffentlicht: (2025)
von: He, Guangzhao, et al.
Veröffentlicht: (2025)
Product of Experts for Visual Generation
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2025)
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2025)
Weakly-Supervised Learning of Dense Functional Correspondences
von: Stojanov, Stefan, et al.
Veröffentlicht: (2025)
von: Stojanov, Stefan, et al.
Veröffentlicht: (2025)
Web-Scale Collection of Video Data for 4D Animal Reconstruction
von: Zhao, Brian Nlong, et al.
Veröffentlicht: (2025)
von: Zhao, Brian Nlong, et al.
Veröffentlicht: (2025)
Ctrl-VI: Controllable Video Synthesis via Variational Inference
von: Duan, Haoyi, et al.
Veröffentlicht: (2025)
von: Duan, Haoyi, et al.
Veröffentlicht: (2025)
Learning to Think Fast and Slow for Visual Language Models
von: Lin, Chenyu, et al.
Veröffentlicht: (2025)
von: Lin, Chenyu, et al.
Veröffentlicht: (2025)
Flow4R: Unifying 4D Reconstruction and Tracking with Scene Flow
von: Qian, Shenhan, et al.
Veröffentlicht: (2026)
von: Qian, Shenhan, et al.
Veröffentlicht: (2026)
Thinking with Spatial Code for Physical-World Video Reasoning
von: Chen, Jieneng, et al.
Veröffentlicht: (2026)
von: Chen, Jieneng, et al.
Veröffentlicht: (2026)
Naturally Supervised 3D Visual Grounding with Language-Regularized Concept Learners
von: Feng, Chun, et al.
Veröffentlicht: (2024)
von: Feng, Chun, et al.
Veröffentlicht: (2024)
Complementary and Contrastive Learning for Audio-Visual Segmentation
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
von: Gong, Sitong, et al.
Veröffentlicht: (2025)
Farm3D: Learning Articulated 3D Animals by Distilling 2D Diffusion
von: Jakab, Tomas, et al.
Veröffentlicht: (2023)
von: Jakab, Tomas, et al.
Veröffentlicht: (2023)
Grounded Question-Answering in Long Egocentric Videos
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
von: Di, Shangzhe, et al.
Veröffentlicht: (2023)
Coupled Diffusion Sampling for Training-Free Multi-View Image Editing
von: Alzayer, Hadi, et al.
Veröffentlicht: (2025)
von: Alzayer, Hadi, et al.
Veröffentlicht: (2025)
Self-Supervised Learning of Motion Concepts by Optimizing Counterfactuals
von: Stojanov, Stefan, et al.
Veröffentlicht: (2025)
von: Stojanov, Stefan, et al.
Veröffentlicht: (2025)
SK-Adapter: Skeleton-Based Structural Control for Native 3D Generation
von: Wang, Anbang, et al.
Veröffentlicht: (2026)
von: Wang, Anbang, et al.
Veröffentlicht: (2026)
DualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstruction
von: Kaye, Ben, et al.
Veröffentlicht: (2024)
von: Kaye, Ben, et al.
Veröffentlicht: (2024)
OmniStream: Mastering Perception, Reconstruction and Action in Continuous Streams
von: Yan, Yibin, et al.
Veröffentlicht: (2026)
von: Yan, Yibin, et al.
Veröffentlicht: (2026)
3D Congealing: 3D-Aware Image Alignment in the Wild
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2024)
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2024)
Look Again, Think Slowly: Enhancing Visual Reflection in Vision-Language Models
von: Jian, Pu, et al.
Veröffentlicht: (2025)
von: Jian, Pu, et al.
Veröffentlicht: (2025)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
Per-Query Visual Concept Learning
von: Malca, Ori, et al.
Veröffentlicht: (2025)
von: Malca, Ori, et al.
Veröffentlicht: (2025)
Cascade Prompt Learning for Vision-Language Model Adaptation
von: Wu, Ge, et al.
Veröffentlicht: (2024)
von: Wu, Ge, et al.
Veröffentlicht: (2024)
Few-Shot Learning of Visual Compositional Concepts through Probabilistic Schema Induction
von: Lee, Andrew Jun, et al.
Veröffentlicht: (2025)
von: Lee, Andrew Jun, et al.
Veröffentlicht: (2025)
Exploring Part-Informed Visual-Language Learning for Person Re-Identification
von: Lin, Yin, et al.
Veröffentlicht: (2023)
von: Lin, Yin, et al.
Veröffentlicht: (2023)
Visual Superordinate Abstraction for Robust Concept Learning
von: Zheng, Qi, et al.
Veröffentlicht: (2022)
von: Zheng, Qi, et al.
Veröffentlicht: (2022)
Visual-Friendly Concept Protection via Selective Adversarial Perturbations
von: Mi, Xiaoyue, et al.
Veröffentlicht: (2024)
von: Mi, Xiaoyue, et al.
Veröffentlicht: (2024)
Boosting Continual Learning of Vision-Language Models via Mixture-of-Experts Adapters
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024)
von: Yu, Jiazuo, et al.
Veröffentlicht: (2024)
Universal Video Temporal Grounding with Generative Multi-modal Large Language Models
von: Li, Zeqian, et al.
Veröffentlicht: (2025)
von: Li, Zeqian, et al.
Veröffentlicht: (2025)
Neuro-Symbolic Concepts
von: Mao, Jiayuan, et al.
Veröffentlicht: (2025)
von: Mao, Jiayuan, et al.
Veröffentlicht: (2025)
Concept-Guided Prompt Learning for Generalization in Vision-Language Models
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
von: Zhang, Yi, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Seeing a Rose in Five Thousand Ways
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2022) -
The Scene Language: Representing Scenes with Programs, Words, and Embeddings
von: Zhang, Yunzhi, et al.
Veröffentlicht: (2024) -
Ponymation: Learning Articulated 3D Animal Motions from Unlabeled Online Videos
von: Sun, Keqiang, et al.
Veröffentlicht: (2023) -
Birth and Death of a Rose
von: Geng, Chen, et al.
Veröffentlicht: (2024) -
NeuROK: Generative 4D Neural Object Kinematics
von: Geng, Chen, et al.
Veröffentlicht: (2026)