Towards More Unified In-context Visual Understanding
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Sheng, Dianmo, Chen, Dongdong, Tan, Zhentao, Liu, Qiankun, Chu, Qi, Bao, Jianmin, Gong, Tao, Liu, Bin, Xu, Shengwei, Yu, Nenghai |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs
von: Zhao, Xuanpu, et al.
Veröffentlicht: (2026)
von: Zhao, Xuanpu, et al.
Veröffentlicht: (2026)
Transformer based Pluralistic Image Completion with Reduced Information Loss
von: Liu, Qiankun, et al.
Veröffentlicht: (2024)
von: Liu, Qiankun, et al.
Veröffentlicht: (2024)
Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
TCI-Former: Thermal Conduction-Inspired Transformer for Infrared Small Target Detection
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
SAPL: Semantic-Agnostic Prompt Learning in CLIP for Weakly Supervised Image Manipulation Localization
von: Wang, Xinghao, et al.
Veröffentlicht: (2026)
von: Wang, Xinghao, et al.
Veröffentlicht: (2026)
MiM-ISTD: Mamba-in-Mamba for Efficient Infrared Small Target Detection
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
Context-Aware Weakly Supervised Image Manipulation Localization with SAM Refinement
von: Wang, Xinghao, et al.
Veröffentlicht: (2025)
von: Wang, Xinghao, et al.
Veröffentlicht: (2025)
FishBEV: Distortion-Resilient Bird's Eye View Segmentation with Surround-View Fisheye Cameras
von: Li, Hang, et al.
Veröffentlicht: (2025)
von: Li, Hang, et al.
Veröffentlicht: (2025)
Multi-spectral Class Center Network for Face Manipulation Detection and Localization
von: Miao, Changtao, et al.
Veröffentlicht: (2023)
von: Miao, Changtao, et al.
Veröffentlicht: (2023)
LAKAN: Landmark-assisted Adaptive Kolmogorov-Arnold Network for Face Forgery Detection
von: Jiang, Jiayao, et al.
Veröffentlicht: (2025)
von: Jiang, Jiayao, et al.
Veröffentlicht: (2025)
Mixture-of-Noises Enhanced Forgery-Aware Predictor for Multi-Face Manipulation Detection and Localization
von: Miao, Changtao, et al.
Veröffentlicht: (2024)
von: Miao, Changtao, et al.
Veröffentlicht: (2024)
Towards Anytime Retrieval: A Benchmark for Anytime Person Re-Identification
von: Li, Xulin, et al.
Veröffentlicht: (2025)
von: Li, Xulin, et al.
Veröffentlicht: (2025)
Llama SLayer 8B: Shallow Layers Hold the Key to Knowledge Injection
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024)
AnyPattern: Towards In-context Image Copy Detection
von: Wang, Wenhao, et al.
Veröffentlicht: (2024)
von: Wang, Wenhao, et al.
Veröffentlicht: (2024)
Training-Free In-Context Forensic Chain for Image Manipulation Detection and Localization
von: Chen, Rui, et al.
Veröffentlicht: (2025)
von: Chen, Rui, et al.
Veröffentlicht: (2025)
Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
von: Wu, Size, et al.
Veröffentlicht: (2025)
von: Wu, Size, et al.
Veröffentlicht: (2025)
Exploiting Modality-Specific Features For Multi-Modal Manipulation Detection And Grounding
von: Wang, Jiazhen, et al.
Veröffentlicht: (2023)
von: Wang, Jiazhen, et al.
Veröffentlicht: (2023)
Advancing Aesthetic Image Generation via Composition Transfer
von: Zou, Kai, et al.
Veröffentlicht: (2026)
von: Zou, Kai, et al.
Veröffentlicht: (2026)
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
von: Song, Wei, et al.
Veröffentlicht: (2025)
von: Song, Wei, et al.
Veröffentlicht: (2025)
XYZCylinder: Towards Compatible Feed-Forward 3D Gaussian Splatting for Driving Scenes via Unified Cylinder Lifting Method
von: Yu, Haochen, et al.
Veröffentlicht: (2025)
von: Yu, Haochen, et al.
Veröffentlicht: (2025)
Flora: Effortless Context Construction to Arbitrary Length and Scale
von: Chen, Tianxiang, et al.
Veröffentlicht: (2025)
von: Chen, Tianxiang, et al.
Veröffentlicht: (2025)
Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling
von: Zhou, Chao, et al.
Veröffentlicht: (2025)
von: Zhou, Chao, et al.
Veröffentlicht: (2025)
MMGen: Unified Multi-modal Image Generation and Understanding in One Go
von: Wang, Jiepeng, et al.
Veröffentlicht: (2025)
von: Wang, Jiepeng, et al.
Veröffentlicht: (2025)
GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
von: Xiang, Yuxiao, et al.
Veröffentlicht: (2025)
von: Xiang, Yuxiao, et al.
Veröffentlicht: (2025)
MedHorizon: Towards Long-context Medical Video Understanding in the Wild
von: Du, Bodong, et al.
Veröffentlicht: (2026)
von: Du, Bodong, et al.
Veröffentlicht: (2026)
Counterfactual Intervention Feature Transfer for Visible-Infrared Person Re-identification
von: Li, Xulin, et al.
Veröffentlicht: (2022)
von: Li, Xulin, et al.
Veröffentlicht: (2022)
HiAR: Efficient Autoregressive Long Video Generation via Hierarchical Denoising
von: Zou, Kai, et al.
Veröffentlicht: (2026)
von: Zou, Kai, et al.
Veröffentlicht: (2026)
UniPercept: Towards Unified Perceptual-Level Image Understanding across Aesthetics, Quality, Structure, and Texture
von: Cao, Shuo, et al.
Veröffentlicht: (2025)
von: Cao, Shuo, et al.
Veröffentlicht: (2025)
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
von: Wu, Size, et al.
Veröffentlicht: (2025)
von: Wu, Size, et al.
Veröffentlicht: (2025)
UniTok: A Unified Tokenizer for Visual Generation and Understanding
von: Ma, Chuofan, et al.
Veröffentlicht: (2025)
von: Ma, Chuofan, et al.
Veröffentlicht: (2025)
StoryImager: A Unified and Efficient Framework for Coherent Story Visualization and Completion
von: Tao, Ming, et al.
Veröffentlicht: (2024)
von: Tao, Ming, et al.
Veröffentlicht: (2024)
UniEmo: Unifying Emotional Understanding and Generation with Learnable Expert Queries
von: Zhu, Yijie, et al.
Veröffentlicht: (2025)
von: Zhu, Yijie, et al.
Veröffentlicht: (2025)
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
von: Wu, Chengyue, et al.
Veröffentlicht: (2024)
Uni-ViGU: Towards Unified Video Generation and Understanding via A Diffusion-Based Video Generator
von: Qin, Luozheng, et al.
Veröffentlicht: (2026)
von: Qin, Luozheng, et al.
Veröffentlicht: (2026)
Omni-Video: Democratizing Unified Video Understanding and Generation
von: Tan, Zhiyu, et al.
Veröffentlicht: (2025)
von: Tan, Zhiyu, et al.
Veröffentlicht: (2025)
MAFE R-CNN: Selecting More Samples to Learn Category-aware Features for Small Object Detection
von: Li, Yichen, et al.
Veröffentlicht: (2025)
von: Li, Yichen, et al.
Veröffentlicht: (2025)
Textured-GS: Gaussian Splatting with Spatially Defined Color and Opacity
von: Huang, Zhentao, et al.
Veröffentlicht: (2024)
von: Huang, Zhentao, et al.
Veröffentlicht: (2024)
OracleSage: Towards Unified Visual-Linguistic Understanding of Oracle Bone Scripts through Cross-Modal Knowledge Fusion
von: Jiang, Hanqi, et al.
Veröffentlicht: (2024)
von: Jiang, Hanqi, et al.
Veröffentlicht: (2024)
TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
von: Zhao, Weichao, et al.
Veröffentlicht: (2024)
MMLongBench-Doc: Benchmarking Long-context Document Understanding with Visualizations
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
von: Ma, Yubo, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs
von: Zhao, Xuanpu, et al.
Veröffentlicht: (2026) -
Transformer based Pluralistic Image Completion with Reduced Information Loss
von: Liu, Qiankun, et al.
Veröffentlicht: (2024) -
Bootstrapping Audio-Visual Segmentation by Strengthening Audio Cues
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024) -
TCI-Former: Thermal Conduction-Inspired Transformer for Infrared Small Target Detection
von: Chen, Tianxiang, et al.
Veröffentlicht: (2024) -
SAPL: Semantic-Agnostic Prompt Learning in CLIP for Weakly Supervised Image Manipulation Localization
von: Wang, Xinghao, et al.
Veröffentlicht: (2026)