Exploring Effective Factors for Improving Visual In-Context Learning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sun, Yanpeng, Chen, Qiang, Li, Xiaofan, Wang, Jian, Wang, Jingdong, Li, Zechao |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
VRP-SAM: SAM with Visual Reference Prompt
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024)
Artemis: Structured Visual Reasoning for Perception Policy Learning
von: Tang, Wei, et al.
Veröffentlicht: (2025)
von: Tang, Wei, et al.
Veröffentlicht: (2025)
FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts
von: Bo, Weihao, et al.
Veröffentlicht: (2025)
von: Bo, Weihao, et al.
Veröffentlicht: (2025)
Visual Position Prompt for MLLM based Visual Grounding
von: Tang, Wei, et al.
Veröffentlicht: (2025)
von: Tang, Wei, et al.
Veröffentlicht: (2025)
SSP-SAM: SAM with Semantic-Spatial Prompt for Referring Expression Segmentation
von: Tang, Wei, et al.
Veröffentlicht: (2026)
von: Tang, Wei, et al.
Veröffentlicht: (2026)
CSGO: Content-Style Composition in Text-to-Image Generation
von: Xing, Peng, et al.
Veröffentlicht: (2024)
von: Xing, Peng, et al.
Veröffentlicht: (2024)
FullAnno: A Data Engine for Enhancing Image Comprehension of MLLMs
von: Hao, Jing, et al.
Veröffentlicht: (2024)
von: Hao, Jing, et al.
Veröffentlicht: (2024)
Improving Generalized Visual Grounding with Instance-aware Joint Learning
von: Dai, Ming, et al.
Veröffentlicht: (2025)
von: Dai, Ming, et al.
Veröffentlicht: (2025)
Pyramidal Patchification Flow for Visual Generation
von: Li, Hui, et al.
Veröffentlicht: (2025)
von: Li, Hui, et al.
Veröffentlicht: (2025)
FVAR: Visual Autoregressive Modeling via Next Focus Prediction
von: Li, Xiaofan, et al.
Veröffentlicht: (2025)
von: Li, Xiaofan, et al.
Veröffentlicht: (2025)
Diff-MM: Exploring Pre-trained Text-to-Image Generation Model for Unified Multi-modal Object Tracking
von: Xuan, Shiyu, et al.
Veröffentlicht: (2025)
von: Xuan, Shiyu, et al.
Veröffentlicht: (2025)
Improving Data Augmentation for Robust Visual Question Answering with Effective Curriculum Learning
von: Zheng, Yuhang, et al.
Veröffentlicht: (2024)
von: Zheng, Yuhang, et al.
Veröffentlicht: (2024)
Continual SFT Matches Multimodal RLHF with Negative Supervision
von: Zhu, Ke, et al.
Veröffentlicht: (2024)
von: Zhu, Ke, et al.
Veröffentlicht: (2024)
Learning Contrastive Self-Distillation for Ultra-Fine-Grained Visual Categorization Targeting Limited Samples
von: Fang, Ziye, et al.
Veröffentlicht: (2023)
von: Fang, Ziye, et al.
Veröffentlicht: (2023)
A Comprehensive Survey on Visual Concept Mining in Text-to-image Diffusion Models
von: Li, Ziqiang, et al.
Veröffentlicht: (2025)
von: Li, Ziqiang, et al.
Veröffentlicht: (2025)
Perception Before Reasoning: Two-Stage Reinforcement Learning for Visual Reasoning in Vision-Language Models
von: Chen, Yan, et al.
Veröffentlicht: (2025)
von: Chen, Yan, et al.
Veröffentlicht: (2025)
Video4Edit: Viewing Image Editing as a Degenerate Temporal Process
von: Li, Xiaofan, et al.
Veröffentlicht: (2025)
von: Li, Xiaofan, et al.
Veröffentlicht: (2025)
VoxelSplat: Dynamic Gaussian Splatting as an Effective Loss for Occupancy and Flow Prediction
von: Zhu, Ziyue, et al.
Veröffentlicht: (2025)
von: Zhu, Ziyue, et al.
Veröffentlicht: (2025)
See the Text: From Tokenization to Visual Reading
von: Xing, Ling, et al.
Veröffentlicht: (2025)
von: Xing, Ling, et al.
Veröffentlicht: (2025)
HumanSense: From Multimodal Perception to Empathetic Context-Aware Responses through Reasoning MLLMs
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
von: Qin, Zheng, et al.
Veröffentlicht: (2025)
Guiding Visual Autoregressive Models through Spectrum Weakening
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
von: Wang, Chaoyang, et al.
Veröffentlicht: (2025)
CoopTrack: Exploring End-to-End Learning for Efficient Cooperative Sequential Perception
von: Zhong, Jiaru, et al.
Veröffentlicht: (2025)
von: Zhong, Jiaru, et al.
Veröffentlicht: (2025)
Revisiting MLLMs: An In-Depth Analysis of Image Classification Abilities
von: Liu, Huan, et al.
Veröffentlicht: (2024)
von: Liu, Huan, et al.
Veröffentlicht: (2024)
Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search
von: Zhang, Jingdong, et al.
Veröffentlicht: (2026)
von: Zhang, Jingdong, et al.
Veröffentlicht: (2026)
Leveraging Visual Tokens for Extended Text Contexts in Multi-Modal Learning
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
von: Wang, Alex Jinpeng, et al.
Veröffentlicht: (2024)
Beyond Single-Sample: Reliable Multi-Sample Distillation for Video Understanding
von: Li, Songlin, et al.
Veröffentlicht: (2026)
von: Li, Songlin, et al.
Veröffentlicht: (2026)
SADL: An Effective In-Context Learning Method for Compositional Visual QA
von: Dang, Long Hoang, et al.
Veröffentlicht: (2024)
von: Dang, Long Hoang, et al.
Veröffentlicht: (2024)
Enhancing DETRs Variants through Improved Content Query and Similar Query Aggregation
von: Zhang, Yingying, et al.
Veröffentlicht: (2024)
von: Zhang, Yingying, et al.
Veröffentlicht: (2024)
MS-DETR: Efficient DETR Training with Mixed Supervision
von: Zhao, Chuyang, et al.
Veröffentlicht: (2024)
von: Zhao, Chuyang, et al.
Veröffentlicht: (2024)
Inv-Adapter: ID Customization Generation via Image Inversion and Lightweight Adapter
von: Xing, Peng, et al.
Veröffentlicht: (2024)
von: Xing, Peng, et al.
Veröffentlicht: (2024)
Zero-shot HOI Detection with MLLM-based Detector-agnostic Interaction Recognition
von: Xuan, Shiyu, et al.
Veröffentlicht: (2026)
von: Xuan, Shiyu, et al.
Veröffentlicht: (2026)
Visual In-Context Learning for Large Vision-Language Models
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
von: Zhou, Yucheng, et al.
Veröffentlicht: (2024)
DreamVAR: Taming Reinforced Visual Autoregressive Model for High-Fidelity Subject-Driven Image Generation
von: Jiang, Xin, et al.
Veröffentlicht: (2026)
von: Jiang, Xin, et al.
Veröffentlicht: (2026)
SEED: A Simple and Effective 3D DETR in Point Clouds
von: Liu, Zhe, et al.
Veröffentlicht: (2024)
von: Liu, Zhe, et al.
Veröffentlicht: (2024)
When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion Models
von: Sun, Zhengyang, et al.
Veröffentlicht: (2026)
von: Sun, Zhengyang, et al.
Veröffentlicht: (2026)
GP-NeRF: Generalized Perception NeRF for Context-Aware 3D Scene Understanding
von: Li, Hao, et al.
Veröffentlicht: (2023)
von: Li, Hao, et al.
Veröffentlicht: (2023)
E-InMeMo: Enhanced Prompting for Visual In-Context Learning
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
von: Zhang, Jiahao, et al.
Veröffentlicht: (2025)
OmniFashion: Towards Generalist Fashion Intelligence via Multi-Task Vision-Language Learning
von: Yang, Zhengwei, et al.
Veröffentlicht: (2026)
von: Yang, Zhengwei, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
VRP-SAM: SAM with Visual Reference Prompt
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024) -
Enhancing Descriptive Captions with Visual Attributes for Multimodal Perception
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024) -
Improving Multi-modal Large Language Model through Boosting Vision Capabilities
von: Sun, Yanpeng, et al.
Veröffentlicht: (2024) -
Artemis: Structured Visual Reasoning for Perception Policy Learning
von: Tang, Wei, et al.
Veröffentlicht: (2025) -
FedMGP: Personalized Federated Learning with Multi-Group Text-Visual Prompts
von: Bo, Weihao, et al.
Veröffentlicht: (2025)