In-Context Translation: Towards Unifying Image Recognition, Processing, and Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Xue, Han, Sun, Qianru, Song, Li, Zhang, Wenjun, Huang, Zhiwu |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Latency-aware Unified Dynamic Networks for Efficient Image Recognition
von: Han, Yizeng, et al.
Veröffentlicht: (2023)
von: Han, Yizeng, et al.
Veröffentlicht: (2023)
Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
von: Pan, Kaihang, et al.
Veröffentlicht: (2024)
Class Is Invariant to Context and Vice Versa: On Learning Invariance for Out-Of-Distribution Generalization
von: Qi, Jiaxin, et al.
Veröffentlicht: (2022)
von: Qi, Jiaxin, et al.
Veröffentlicht: (2022)
Weakly-Supervised Semantic Segmentation with Image-Level Labels: from Traditional Models to Foundation Models
von: Chen, Zhaozheng, et al.
Veröffentlicht: (2023)
von: Chen, Zhaozheng, et al.
Veröffentlicht: (2023)
OpenSDI: Spotting Diffusion-Generated Images in the Open World
von: Wang, Yabin, et al.
Veröffentlicht: (2025)
von: Wang, Yabin, et al.
Veröffentlicht: (2025)
ADDP: Learning General Representations for Image Recognition and Generation with Alternating Denoising Diffusion Process
von: Tian, Changyao, et al.
Veröffentlicht: (2023)
von: Tian, Changyao, et al.
Veröffentlicht: (2023)
Towards Natural Image Matting in the Wild via Real-Scenario Prior
von: Xia, Ruihao, et al.
Veröffentlicht: (2024)
von: Xia, Ruihao, et al.
Veröffentlicht: (2024)
EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
von: Ju, Xuan, et al.
Veröffentlicht: (2025)
von: Ju, Xuan, et al.
Veröffentlicht: (2025)
MICON-Bench: Benchmarking and Enhancing Multi-Image Context Image Generation in Unified Multimodal Models
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
von: Wu, Mingrui, et al.
Veröffentlicht: (2026)
Global-Local Detail Guided Transformer for Sea Ice Recognition in Optical Remote Sensing Images
von: Huang, Zhanchao, et al.
Veröffentlicht: (2024)
von: Huang, Zhanchao, et al.
Veröffentlicht: (2024)
Query-Kontext: An Unified Multimodal Model for Image Generation and Editing
von: Song, Yuxin, et al.
Veröffentlicht: (2025)
von: Song, Yuxin, et al.
Veröffentlicht: (2025)
Towards Unified Facial Action Unit Recognition Framework by Large Language Models
von: Hu, Guohong, et al.
Veröffentlicht: (2024)
von: Hu, Guohong, et al.
Veröffentlicht: (2024)
Generalized Logit Adjustment: Calibrating Fine-tuned Models by Removing Label Bias in Foundation Models
von: Zhu, Beier, et al.
Veröffentlicht: (2023)
von: Zhu, Beier, et al.
Veröffentlicht: (2023)
Video Anomaly Detection and Explanation via Large Language Models
von: Lv, Hui, et al.
Veröffentlicht: (2024)
von: Lv, Hui, et al.
Veröffentlicht: (2024)
Target Recognition Algorithm for Monitoring Images in Electric Power Construction Process
von: Song, Hao, et al.
Veröffentlicht: (2024)
von: Song, Hao, et al.
Veröffentlicht: (2024)
Zero-shot Referring Expression Comprehension via Structural Similarity Between Images and Captions
von: Han, Zeyu, et al.
Veröffentlicht: (2023)
von: Han, Zeyu, et al.
Veröffentlicht: (2023)
Medical Vision Generalist: Unifying Medical Imaging Tasks in Context
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
von: Ren, Sucheng, et al.
Veröffentlicht: (2024)
UNIT: Unifying Image and Text Recognition in One Vision Encoder
von: Zhu, Yi, et al.
Veröffentlicht: (2024)
von: Zhu, Yi, et al.
Veröffentlicht: (2024)
Unified Generative and Discriminative Training for Multi-modal Large Language Models
von: Chow, Wei, et al.
Veröffentlicht: (2024)
von: Chow, Wei, et al.
Veröffentlicht: (2024)
Towards Generalized Multi-Image Editing for Unified Multimodal Models
von: Xu, Pengcheng, et al.
Veröffentlicht: (2026)
von: Xu, Pengcheng, et al.
Veröffentlicht: (2026)
Linguistic Profiling of Deepfakes: An Open Database for Next-Generation Deepfake Detection
von: Wang, Yabin, et al.
Veröffentlicht: (2024)
von: Wang, Yabin, et al.
Veröffentlicht: (2024)
UNIMO-G: Unified Image Generation through Multimodal Conditional Diffusion
von: Li, Wei, et al.
Veröffentlicht: (2024)
von: Li, Wei, et al.
Veröffentlicht: (2024)
Towards Enhanced Image Generation Via Multi-modal Chain of Thought in Unified Generative Models
von: Wang, Yi, et al.
Veröffentlicht: (2025)
von: Wang, Yi, et al.
Veröffentlicht: (2025)
Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
von: Nie, Ming, et al.
Veröffentlicht: (2026)
von: Nie, Ming, et al.
Veröffentlicht: (2026)
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation
von: Zhang, Yabo, et al.
Veröffentlicht: (2026)
von: Zhang, Yabo, et al.
Veröffentlicht: (2026)
UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation
von: Duan, Lunhao, et al.
Veröffentlicht: (2024)
von: Duan, Lunhao, et al.
Veröffentlicht: (2024)
Towards Online Continuous Sign Language Recognition and Translation
von: Zuo, Ronglai, et al.
Veröffentlicht: (2024)
von: Zuo, Ronglai, et al.
Veröffentlicht: (2024)
UniDGF: A Unified Detection-to-Generation Framework for Hierarchical Object Visual Recognition
von: Nan, Xinyu, et al.
Veröffentlicht: (2025)
von: Nan, Xinyu, et al.
Veröffentlicht: (2025)
CoTMR: Chain-of-Thought Multi-Scale Reasoning for Training-Free Zero-Shot Composed Image Retrieval
von: Sun, Zelong, et al.
Veröffentlicht: (2025)
von: Sun, Zelong, et al.
Veröffentlicht: (2025)
PRIM: Towards Practical In-Image Multilingual Machine Translation
von: Tian, Yanzhi, et al.
Veröffentlicht: (2025)
von: Tian, Yanzhi, et al.
Veröffentlicht: (2025)
Region-Wise Correspondence Prediction between Manga Line Art Images
von: Li, Yingxuan, et al.
Veröffentlicht: (2025)
von: Li, Yingxuan, et al.
Veröffentlicht: (2025)
UniAlignment: Semantic Alignment for Unified Image Generation, Understanding, Manipulation and Perception
von: Song, Xinyang, et al.
Veröffentlicht: (2025)
von: Song, Xinyang, et al.
Veröffentlicht: (2025)
3SGen: Unified Subject, Style, and Structure-Driven Image Generation with Adaptive Task-specific Memory
von: Song, Xinyang, et al.
Veröffentlicht: (2025)
von: Song, Xinyang, et al.
Veröffentlicht: (2025)
UniFormer: Unifying Convolution and Self-attention for Visual Recognition
von: Li, Kunchang, et al.
Veröffentlicht: (2022)
von: Li, Kunchang, et al.
Veröffentlicht: (2022)
A Cosine Network for Image Super-Resolution
von: Tian, Chunwei, et al.
Veröffentlicht: (2026)
von: Tian, Chunwei, et al.
Veröffentlicht: (2026)
Unified Restoration-Perception Learning: Maritime Infrared-Visible Image Fusion and Segmentation
von: Cai, Weichao, et al.
Veröffentlicht: (2026)
von: Cai, Weichao, et al.
Veröffentlicht: (2026)
Generalized Visual Relation Detection with Diffusion Models
von: Gao, Kaifeng, et al.
Veröffentlicht: (2025)
von: Gao, Kaifeng, et al.
Veröffentlicht: (2025)
Dyn-Adapter: Towards Disentangled Representation for Efficient Visual Recognition
von: Zhang, Yurong, et al.
Veröffentlicht: (2024)
von: Zhang, Yurong, et al.
Veröffentlicht: (2024)
Degradation-Aware Adaptive Context Gating for Unified Image Restoration
von: He, Lei, et al.
Veröffentlicht: (2026)
von: He, Lei, et al.
Veröffentlicht: (2026)
GenN2N: Generative NeRF2NeRF Translation
von: Liu, Xiangyue, et al.
Veröffentlicht: (2024)
von: Liu, Xiangyue, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Latency-aware Unified Dynamic Networks for Efficient Image Recognition
von: Han, Yizeng, et al.
Veröffentlicht: (2023) -
Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration
von: Pan, Kaihang, et al.
Veröffentlicht: (2024) -
Class Is Invariant to Context and Vice Versa: On Learning Invariance for Out-Of-Distribution Generalization
von: Qi, Jiaxin, et al.
Veröffentlicht: (2022) -
Weakly-Supervised Semantic Segmentation with Image-Level Labels: from Traditional Models to Foundation Models
von: Chen, Zhaozheng, et al.
Veröffentlicht: (2023) -
OpenSDI: Spotting Diffusion-Generated Images in the Open World
von: Wang, Yabin, et al.
Veröffentlicht: (2025)