Visual Transformation Telling
Fuente:
arXiv
Salvato in:
| Autori principali: | Cui, Wanqing, Hong, Xin, Lan, Yanyan, Pang, Liang, Guo, Jiafeng, Cheng, Xueqi |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2023
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
di: Cui, Wanqing, et al.
Pubblicazione: (2024)
di: Cui, Wanqing, et al.
Pubblicazione: (2024)
Classifier Guidance Enhances Diffusion-based Adversarial Purification by Preserving Predictive Information
di: Zhang, Mingkun, et al.
Pubblicazione: (2024)
di: Zhang, Mingkun, et al.
Pubblicazione: (2024)
CausalDiff: Causality-Inspired Disentanglement via Diffusion Model for Adversarial Defense
di: Zhang, Mingkun, et al.
Pubblicazione: (2024)
di: Zhang, Mingkun, et al.
Pubblicazione: (2024)
CLIPure: Purification in Latent Space via CLIP for Adversarially Robust Zero-Shot Classification
di: Zhang, Mingkun, et al.
Pubblicazione: (2025)
di: Zhang, Mingkun, et al.
Pubblicazione: (2025)
Improving Video Corpus Moment Retrieval with Partial Relevance Enhancement
di: Hou, Danyang, et al.
Pubblicazione: (2024)
di: Hou, Danyang, et al.
Pubblicazione: (2024)
Event-aware Video Corpus Moment Retrieval
di: Hou, Danyang, et al.
Pubblicazione: (2024)
di: Hou, Danyang, et al.
Pubblicazione: (2024)
Focusing by Contrastive Attention: Enhancing VLMs' Visual Reasoning
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
di: Ge, Yuyao, et al.
Pubblicazione: (2025)
Compression Tells Intelligence: Visual Coding, Visual Token Technology, and the Unification
di: Jin, Xin, et al.
Pubblicazione: (2026)
di: Jin, Xin, et al.
Pubblicazione: (2026)
HighlightBench: Benchmarking Markup-Driven Table Reasoning in Scientific Documents
di: Wang, Lexin, et al.
Pubblicazione: (2026)
di: Wang, Lexin, et al.
Pubblicazione: (2026)
Cross-Modal Safety Mechanism Transfer in Large Vision-Language Models
di: Xu, Shicheng, et al.
Pubblicazione: (2024)
di: Xu, Shicheng, et al.
Pubblicazione: (2024)
Tell Me Without Telling Me: Two-Way Prediction of Visualization Literacy and Visual Attention
di: Chang, Minsuk, et al.
Pubblicazione: (2025)
di: Chang, Minsuk, et al.
Pubblicazione: (2025)
Attention Grounded Enhancement for Visual Document Retrieval
di: Cui, Wanqing, et al.
Pubblicazione: (2025)
di: Cui, Wanqing, et al.
Pubblicazione: (2025)
Transformer-Based Visual Segmentation: A Survey
di: Li, Xiangtai, et al.
Pubblicazione: (2023)
di: Li, Xiangtai, et al.
Pubblicazione: (2023)
CARE Transformer: Mobile-Friendly Linear Visual Transformer via Decoupled Dual Interaction
di: Zhou, Yuan, et al.
Pubblicazione: (2024)
di: Zhou, Yuan, et al.
Pubblicazione: (2024)
4DLangVGGT: 4D Language-Visual Geometry Grounded Transformer
di: Wu, Xianfeng, et al.
Pubblicazione: (2025)
di: Wu, Xianfeng, et al.
Pubblicazione: (2025)
Show or Tell? A Benchmark To Evaluate Visual and Textual Prompts in Semantic Segmentation
di: Rosi, Gabriele, et al.
Pubblicazione: (2025)
di: Rosi, Gabriele, et al.
Pubblicazione: (2025)
Fast Multi-Organ Fine Segmentation in CT Images with Hierarchical Sparse Sampling and Residual Transformer
di: Guo, Xueqi, et al.
Pubblicazione: (2025)
di: Guo, Xueqi, et al.
Pubblicazione: (2025)
FeTT: Continual Class Incremental Learning via Feature Transformation Tuning
di: Qiang, Sunyuan, et al.
Pubblicazione: (2024)
di: Qiang, Sunyuan, et al.
Pubblicazione: (2024)
MangaDiT: Reference-Guided Line Art Colorization with Hierarchical Attention in Diffusion Transformers
di: Qiu, Qianru, et al.
Pubblicazione: (2025)
di: Qiu, Qianru, et al.
Pubblicazione: (2025)
Aligning Generative Denoising with Discriminative Objectives Unleashes Diffusion for Visual Perception
di: Pang, Ziqi, et al.
Pubblicazione: (2025)
di: Pang, Ziqi, et al.
Pubblicazione: (2025)
LLaVA-SLT: Visual Language Tuning for Sign Language Translation
di: Liang, Han, et al.
Pubblicazione: (2024)
di: Liang, Han, et al.
Pubblicazione: (2024)
Diffusion Is Your Friend in Show, Suggest and Tell
di: Hu, Jia Cheng, et al.
Pubblicazione: (2025)
di: Hu, Jia Cheng, et al.
Pubblicazione: (2025)
Mixture of Balanced Information Bottlenecks for Long-Tailed Visual Recognition
di: Lan, Yifan, et al.
Pubblicazione: (2025)
di: Lan, Yifan, et al.
Pubblicazione: (2025)
Generative Ghost: Investigating Ranking Bias Hidden in AI-Generated Videos
di: Gao, Haowen, et al.
Pubblicazione: (2025)
di: Gao, Haowen, et al.
Pubblicazione: (2025)
Tell Model Where to Look: Mitigating Hallucinations in MLLMs by Vision-Guided Attention
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
di: Zhao, Jianfei, et al.
Pubblicazione: (2025)
EmbodiedPlace: Learning Mixture-of-Features with Embodied Constraints for Visual Place Recognition
di: Liu, Bingxi, et al.
Pubblicazione: (2025)
di: Liu, Bingxi, et al.
Pubblicazione: (2025)
LatentRouter: Can We Choose the Right Multimodal Model Before Seeing Its Answer?
di: Cheng, Xueqi, et al.
Pubblicazione: (2026)
di: Cheng, Xueqi, et al.
Pubblicazione: (2026)
VGGT4D: Mining Motion Cues in Visual Geometry Transformers for 4D Scene Reconstruction
di: Hu, Yu, et al.
Pubblicazione: (2025)
di: Hu, Yu, et al.
Pubblicazione: (2025)
LORE: Latent Optimization for Precise Semantic Control in Rectified Flow-based Image Editing
di: Ouyang, Liangyang, et al.
Pubblicazione: (2025)
di: Ouyang, Liangyang, et al.
Pubblicazione: (2025)
Show and Tell: Visually Explainable Deep Neural Nets via Spatially-Aware Concept Bottleneck Models
di: Benou, Itay, et al.
Pubblicazione: (2025)
di: Benou, Itay, et al.
Pubblicazione: (2025)
Let Storytelling Tell Vivid Stories: An Expressive and Fluent Multimodal Storyteller
di: Zang, Chuanqi, et al.
Pubblicazione: (2024)
di: Zang, Chuanqi, et al.
Pubblicazione: (2024)
HD-VGGT: High-Resolution Visual Geometry Transformer
di: Chen, Tianrun, et al.
Pubblicazione: (2026)
di: Chen, Tianrun, et al.
Pubblicazione: (2026)
Show, Don't Tell: Morphing Latent Reasoning into Image Generation
di: Chen, Harold Haodong, et al.
Pubblicazione: (2026)
di: Chen, Harold Haodong, et al.
Pubblicazione: (2026)
Unifying Feature and Cost Aggregation with Transformers for Semantic and Visual Correspondence
di: Hong, Sunghwan, et al.
Pubblicazione: (2024)
di: Hong, Sunghwan, et al.
Pubblicazione: (2024)
VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
di: Yang, Sihan, et al.
Pubblicazione: (2025)
di: Yang, Sihan, et al.
Pubblicazione: (2025)
big.LITTLE Vision Transformer for Efficient Visual Recognition
di: Guo, He, et al.
Pubblicazione: (2024)
di: Guo, He, et al.
Pubblicazione: (2024)
LVIC: Multi-modality segmentation by Lifting Visual Info as Cue
di: Dong, Zichao, et al.
Pubblicazione: (2024)
di: Dong, Zichao, et al.
Pubblicazione: (2024)
UGround: Towards Unified Visual Grounding with Unrolled Transformers
di: Qian, Rui, et al.
Pubblicazione: (2025)
di: Qian, Rui, et al.
Pubblicazione: (2025)
Tell Codec What Worth Compressing: Semantically Disentangled Image Coding for Machine with LMMs
di: Liu, Jinming, et al.
Pubblicazione: (2024)
di: Liu, Jinming, et al.
Pubblicazione: (2024)
BSViT: A Burst Spiking Vision Transformer for Expressive and Efficient Visual Representation Learning
di: Peng, Hongxiang, et al.
Pubblicazione: (2026)
di: Peng, Hongxiang, et al.
Pubblicazione: (2026)
Documenti analoghi
-
MVAM: Multi-View Attention Method for Fine-grained Image-Text Matching
di: Cui, Wanqing, et al.
Pubblicazione: (2024) -
Classifier Guidance Enhances Diffusion-based Adversarial Purification by Preserving Predictive Information
di: Zhang, Mingkun, et al.
Pubblicazione: (2024) -
CausalDiff: Causality-Inspired Disentanglement via Diffusion Model for Adversarial Defense
di: Zhang, Mingkun, et al.
Pubblicazione: (2024) -
CLIPure: Purification in Latent Space via CLIP for Adversarially Robust Zero-Shot Classification
di: Zhang, Mingkun, et al.
Pubblicazione: (2025) -
Improving Video Corpus Moment Retrieval with Partial Relevance Enhancement
di: Hou, Danyang, et al.
Pubblicazione: (2024)