Gespeichert in:
| Hauptverfasser: | Salehi, Sogand, Shafiei, Mahdi, Yeo, Teresa, Bachmann, Roman, Zamir, Amir |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2407.17365 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Controlled Training Data Generation with Diffusion Models
von: Yeo, Teresa, et al.
Veröffentlicht: (2024)
von: Yeo, Teresa, et al.
Veröffentlicht: (2024)
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
von: Ramachandran, Rahul, et al.
Veröffentlicht: (2025)
von: Ramachandran, Rahul, et al.
Veröffentlicht: (2025)
ViPO: Visual Preference Optimization at Scale
von: Li, Ming, et al.
Veröffentlicht: (2026)
von: Li, Ming, et al.
Veröffentlicht: (2026)
Deep learning-based Visual Measurement Extraction within an Adaptive Digital Twin Framework from Limited Data Using Transfer Learning
von: Dizaji, Mehrdad Shafiei
Veröffentlicht: (2024)
von: Dizaji, Mehrdad Shafiei
Veröffentlicht: (2024)
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
von: Bachmann, Roman, et al.
Veröffentlicht: (2024)
von: Bachmann, Roman, et al.
Veröffentlicht: (2024)
6D Pose Estimation via Keypoint Heatmap Regression with RGB-D Residual Neural Networks
von: Aljosevic, Ismail, et al.
Veröffentlicht: (2026)
von: Aljosevic, Ismail, et al.
Veröffentlicht: (2026)
Per-Query Visual Concept Learning
von: Malca, Ori, et al.
Veröffentlicht: (2025)
von: Malca, Ori, et al.
Veröffentlicht: (2025)
VideoFlexTok: Flexible-Length Coarse-to-Fine Video Tokenization
von: Atanov, Andrei, et al.
Veröffentlicht: (2026)
von: Atanov, Andrei, et al.
Veröffentlicht: (2026)
FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
von: Bachmann, Roman, et al.
Veröffentlicht: (2025)
von: Bachmann, Roman, et al.
Veröffentlicht: (2025)
(1D) Ordered Tokens Enable Efficient Test-Time Search
von: Gao, Zhitong, et al.
Veröffentlicht: (2026)
von: Gao, Zhitong, et al.
Veröffentlicht: (2026)
Personalizing Text-to-Image Generation to Individual Taste
von: Maerten, Anne-Sofie, et al.
Veröffentlicht: (2026)
von: Maerten, Anne-Sofie, et al.
Veröffentlicht: (2026)
Med-PerSAM: One-Shot Visual Prompt Tuning for Personalized Segment Anything Model in Medical Domain
von: Yoon, Hangyul, et al.
Veröffentlicht: (2024)
von: Yoon, Hangyul, et al.
Veröffentlicht: (2024)
ViKey: Enhancing Temporal Understanding in Videos via Visual Prompting
von: Lee, Yeonkyung, et al.
Veröffentlicht: (2026)
von: Lee, Yeonkyung, et al.
Veröffentlicht: (2026)
HuViDPO:Enhancing Video Generation through Direct Preference Optimization for Human-Centric Alignment
von: Jiang, Lifan, et al.
Veröffentlicht: (2025)
von: Jiang, Lifan, et al.
Veröffentlicht: (2025)
Solving Vision Tasks with Simple Photoreceptors Instead of Cameras
von: Atanov, Andrei, et al.
Veröffentlicht: (2024)
von: Atanov, Andrei, et al.
Veröffentlicht: (2024)
AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization
von: Lu, Jinda, et al.
Veröffentlicht: (2025)
von: Lu, Jinda, et al.
Veröffentlicht: (2025)
PSTF-AttControl: Per-Subject-Tuning-Free Personalized Image Generation with Controllable Face Attributes
von: liu, Xiang, et al.
Veröffentlicht: (2025)
von: liu, Xiang, et al.
Veröffentlicht: (2025)
PersonViT: Large-scale Self-supervised Vision Transformer for Person Re-Identification
von: Hu, Bin, et al.
Veröffentlicht: (2024)
von: Hu, Bin, et al.
Veröffentlicht: (2024)
DesignPref: Capturing Personal Preferences in Visual Design Generation
von: Peng, Yi-Hao, et al.
Veröffentlicht: (2025)
von: Peng, Yi-Hao, et al.
Veröffentlicht: (2025)
PerTouch: VLM-Driven Agent for Personalized and Semantic Image Retouching
von: Chang, Zewei, et al.
Veröffentlicht: (2025)
von: Chang, Zewei, et al.
Veröffentlicht: (2025)
PerSense: Training-Free Personalized Instance Segmentation in Dense Images
von: Siddiqui, Muhammad Ibraheem, et al.
Veröffentlicht: (2024)
von: Siddiqui, Muhammad Ibraheem, et al.
Veröffentlicht: (2024)
LoFA: Learning to Predict Personalized Priors for Fast Adaptation of Visual Generative Models
von: Hao, Yiming, et al.
Veröffentlicht: (2025)
von: Hao, Yiming, et al.
Veröffentlicht: (2025)
One Sentence, One Drama: Personalized Short-Form Drama Generation via Multi-Agent Systems
von: Shi, Yufei, et al.
Veröffentlicht: (2026)
von: Shi, Yufei, et al.
Veröffentlicht: (2026)
Personal Visual Context Learning in Large Multimodal Models
von: Xue, Zihui, et al.
Veröffentlicht: (2026)
von: Xue, Zihui, et al.
Veröffentlicht: (2026)
Learning User Preferences for Image Generation Model
von: Mo, Wenyi, et al.
Veröffentlicht: (2025)
von: Mo, Wenyi, et al.
Veröffentlicht: (2025)
MultiViPerFrOG: A Globally Optimized Multi-Viewpoint Perception Framework for Camera Motion and Tissue Deformation
von: Caccianiga, Guido, et al.
Veröffentlicht: (2024)
von: Caccianiga, Guido, et al.
Veröffentlicht: (2024)
ViUniT: Visual Unit Tests for More Robust Visual Programming
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2024)
von: Panagopoulou, Artemis, et al.
Veröffentlicht: (2024)
ViTCoP: Accelerating Large Vision-Language Models via Visual and Textual Semantic Collaborative Pruning
von: Luo, Wen, et al.
Veröffentlicht: (2026)
von: Luo, Wen, et al.
Veröffentlicht: (2026)
Seeing What Matters: Visual Preference Policy Optimization for Visual Generation
von: Ni, Ziqi, et al.
Veröffentlicht: (2025)
von: Ni, Ziqi, et al.
Veröffentlicht: (2025)
MoViE: Mobile Diffusion for Video Editing
von: Karjauv, Adil, et al.
Veröffentlicht: (2024)
von: Karjauv, Adil, et al.
Veröffentlicht: (2024)
ViCToR: Improving Visual Comprehension via Token Reconstruction for Pretraining LMMs
von: Xie, Yin, et al.
Veröffentlicht: (2024)
von: Xie, Yin, et al.
Veröffentlicht: (2024)
ViSpeak: Visual Instruction Feedback in Streaming Videos
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
von: Fu, Shenghao, et al.
Veröffentlicht: (2025)
LoopViT: Scaling Visual ARC with Looped Transformers
von: Shu, Wen-Jie, et al.
Veröffentlicht: (2026)
von: Shu, Wen-Jie, et al.
Veröffentlicht: (2026)
GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning
von: Wu, Fengyi, et al.
Veröffentlicht: (2025)
von: Wu, Fengyi, et al.
Veröffentlicht: (2025)
PersGuard: Preventing Malicious Personalization via Backdoor Attacks on Pre-trained Text-to-Image Diffusion Models
von: Liu, Xinwei, et al.
Veröffentlicht: (2025)
von: Liu, Xinwei, et al.
Veröffentlicht: (2025)
Personalized Vision via Visual In-Context Learning
von: Jiang, Yuxin, et al.
Veröffentlicht: (2025)
von: Jiang, Yuxin, et al.
Veröffentlicht: (2025)
PersONAL: Towards a Comprehensive Benchmark for Personalized Embodied Agents
von: Ziliotto, Filippo, et al.
Veröffentlicht: (2025)
von: Ziliotto, Filippo, et al.
Veröffentlicht: (2025)
ViGoR-Bench: How Far Are Visual Generative Models From Zero-Shot Visual Reasoners?
von: Han, Haonan, et al.
Veröffentlicht: (2026)
von: Han, Haonan, et al.
Veröffentlicht: (2026)
GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
von: Zhang, Peirong, et al.
Veröffentlicht: (2025)
Can World Simulators Reason? Gen-ViRe: A Generative Visual Reasoning Benchmark
von: Liu, Xinxin, et al.
Veröffentlicht: (2025)
von: Liu, Xinxin, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Controlled Training Data Generation with Diffusion Models
von: Yeo, Teresa, et al.
Veröffentlicht: (2024) -
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
von: Ramachandran, Rahul, et al.
Veröffentlicht: (2025) -
ViPO: Visual Preference Optimization at Scale
von: Li, Ming, et al.
Veröffentlicht: (2026) -
Deep learning-based Visual Measurement Extraction within an Adaptive Digital Twin Framework from Limited Data Using Transfer Learning
von: Dizaji, Mehrdad Shafiei
Veröffentlicht: (2024) -
4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities
von: Bachmann, Roman, et al.
Veröffentlicht: (2024)