Salvato in:
| Autori principali: | Yue, Kaiyu, Jia, Menglin, Hou, Ji, Goldstein, Tom |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | https://arxiv.org/abs/2602.15030 |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Zero-Shot Vision Encoder Grafting via LLM Surrogates
di: Yue, Kaiyu, et al.
Pubblicazione: (2025)
di: Yue, Kaiyu, et al.
Pubblicazione: (2025)
Efficient Image Synthesis with Sphere Latent Encoder
di: Do, Tung, et al.
Pubblicazione: (2026)
di: Do, Tung, et al.
Pubblicazione: (2026)
Object Recognition as Next Token Prediction
di: Yue, Kaiyu, et al.
Pubblicazione: (2023)
di: Yue, Kaiyu, et al.
Pubblicazione: (2023)
From Pixels to Prose: A Large Dataset of Dense Image Captions
di: Singla, Vasu, et al.
Pubblicazione: (2024)
di: Singla, Vasu, et al.
Pubblicazione: (2024)
Language-Image Alignment with Fixed Text Encoders
di: Yang, Jingfeng, et al.
Pubblicazione: (2025)
di: Yang, Jingfeng, et al.
Pubblicazione: (2025)
FlowBypass: Rectified Flow Trajectory Bypass for Training-Free Image Editing
di: Han, Menglin, et al.
Pubblicazione: (2026)
di: Han, Menglin, et al.
Pubblicazione: (2026)
UNIT: Unifying Image and Text Recognition in One Vision Encoder
di: Zhu, Yi, et al.
Pubblicazione: (2024)
di: Zhu, Yi, et al.
Pubblicazione: (2024)
Flow Matching Posterior Sampling: A Training-free Conditional Generation for Flow Matching
di: Song, Kaiyu, et al.
Pubblicazione: (2024)
di: Song, Kaiyu, et al.
Pubblicazione: (2024)
General Vision Encoder Features as Guidance in Medical Image Registration
di: Kögl, Fryderyk, et al.
Pubblicazione: (2024)
di: Kögl, Fryderyk, et al.
Pubblicazione: (2024)
Modality-Aware Bias Mitigation and Invariance Learning for Unsupervised Visible-Infrared Person Re-Identification
di: Wang, Menglin, et al.
Pubblicazione: (2025)
di: Wang, Menglin, et al.
Pubblicazione: (2025)
Prior-Constrained Association Learning for Fine-Grained Generalized Category Discovery
di: Wang, Menglin, et al.
Pubblicazione: (2025)
di: Wang, Menglin, et al.
Pubblicazione: (2025)
Encoder-Only Image Registration
di: Chen, Xiang, et al.
Pubblicazione: (2025)
di: Chen, Xiang, et al.
Pubblicazione: (2025)
Think-Then-Generate: Reasoning-Aware Text-to-Image Diffusion with LLM Encoders
di: Kou, Siqi, et al.
Pubblicazione: (2026)
di: Kou, Siqi, et al.
Pubblicazione: (2026)
FineGRAIN: Evaluating Failure Modes of Text-to-Image Models with Vision Language Model Judges
di: Hayes, Kevin David, et al.
Pubblicazione: (2025)
di: Hayes, Kevin David, et al.
Pubblicazione: (2025)
TextBoost: Boosting Text Encoder for Personalized Text-to-Image Generation
di: Park, NaHyeon, et al.
Pubblicazione: (2024)
di: Park, NaHyeon, et al.
Pubblicazione: (2024)
Semantically Robust Unsupervised Image Translation for Paired Remote Sensing Images
di: Fang, Sheng, et al.
Pubblicazione: (2025)
di: Fang, Sheng, et al.
Pubblicazione: (2025)
Text-Guided Semantic Image Encoder
di: Thirukovalluru, Raghuveer, et al.
Pubblicazione: (2025)
di: Thirukovalluru, Raghuveer, et al.
Pubblicazione: (2025)
IDProtector: An Adversarial Noise Encoder to Protect Against ID-Preserving Image Generation
di: Song, Yiren, et al.
Pubblicazione: (2024)
di: Song, Yiren, et al.
Pubblicazione: (2024)
Vision-Based Localization in Dense Urban Environments: A Case Study of an Urban Village in China
di: Wu, Menglin, et al.
Pubblicazione: (2026)
di: Wu, Menglin, et al.
Pubblicazione: (2026)
MobileDiffusion: Instant Text-to-Image Generation on Mobile Devices
di: Zhao, Yang, et al.
Pubblicazione: (2023)
di: Zhao, Yang, et al.
Pubblicazione: (2023)
PhyGDPO: Physics-Aware Groupwise Direct Preference Optimization for Physically Consistent Text-to-Video Generation
di: Cai, Yuanhao, et al.
Pubblicazione: (2025)
di: Cai, Yuanhao, et al.
Pubblicazione: (2025)
Patch-enhanced Mask Encoder Prompt Image Generation
di: Xu, Shusong, et al.
Pubblicazione: (2024)
di: Xu, Shusong, et al.
Pubblicazione: (2024)
General Purpose Image Encoder DINOv2 for Medical Image Registration
di: Song, Xinrui, et al.
Pubblicazione: (2024)
di: Song, Xinrui, et al.
Pubblicazione: (2024)
Topology Sculptor, Shape Refiner: Discrete Diffusion Model for High-Fidelity 3D Meshes Generation
di: Song, Kaiyu, et al.
Pubblicazione: (2025)
di: Song, Kaiyu, et al.
Pubblicazione: (2025)
Adaptive Caching for Faster Video Generation with Diffusion Transformers
di: Kahatapitiya, Kumara, et al.
Pubblicazione: (2024)
di: Kahatapitiya, Kumara, et al.
Pubblicazione: (2024)
Rethinking Oversaturation in Classifier-Free Guidance via Low Frequency
di: Song, Kaiyu, et al.
Pubblicazione: (2025)
di: Song, Kaiyu, et al.
Pubblicazione: (2025)
Leveraging Previous Steps: A Training-free Fast Solver for Flow Diffusion
di: Song, Kaiyu, et al.
Pubblicazione: (2024)
di: Song, Kaiyu, et al.
Pubblicazione: (2024)
Improving Training-free Conditional Diffusion Model via Fisher Information
di: Song, Kaiyu, et al.
Pubblicazione: (2024)
di: Song, Kaiyu, et al.
Pubblicazione: (2024)
Zebra-CoT: A Dataset for Interleaved Vision Language Reasoning
di: Li, Ang, et al.
Pubblicazione: (2025)
di: Li, Ang, et al.
Pubblicazione: (2025)
Neural Light Spheres for Implicit Image Stitching and View Synthesis
di: Chugunov, Ilya, et al.
Pubblicazione: (2024)
di: Chugunov, Ilya, et al.
Pubblicazione: (2024)
SphereDrag: Spherical Geometry-Aware Panoramic Image Editing
di: Feng, Zhiao, et al.
Pubblicazione: (2025)
di: Feng, Zhiao, et al.
Pubblicazione: (2025)
Analysis of Attention in Video Diffusion Transformers
di: Wen, Yuxin, et al.
Pubblicazione: (2025)
di: Wen, Yuxin, et al.
Pubblicazione: (2025)
ARGUS: Hallucination and Omission Evaluation in Video-LLMs
di: Rawal, Ruchit, et al.
Pubblicazione: (2025)
di: Rawal, Ruchit, et al.
Pubblicazione: (2025)
Both Semantics and Reconstruction Matter: Making Representation Encoders Ready for Text-to-Image Generation and Editing
di: Zhang, Shilong, et al.
Pubblicazione: (2025)
di: Zhang, Shilong, et al.
Pubblicazione: (2025)
Covariance Descriptors Meet General Vision Encoders: Riemannian Deep Learning for Medical Image Classification
di: Mayr, Josef, et al.
Pubblicazione: (2025)
di: Mayr, Josef, et al.
Pubblicazione: (2025)
PromptFusion: Decoupling Stability and Plasticity for Continual Learning
di: Chen, Haoran, et al.
Pubblicazione: (2023)
di: Chen, Haoran, et al.
Pubblicazione: (2023)
Latent Enhancing AutoEncoder for Occluded Image Classification
di: Kotwal, Ketan, et al.
Pubblicazione: (2024)
di: Kotwal, Ketan, et al.
Pubblicazione: (2024)
Breaking the Encoder Barrier for Seamless Video-Language Understanding
di: Li, Handong, et al.
Pubblicazione: (2025)
di: Li, Handong, et al.
Pubblicazione: (2025)
Video Prediction Models as General Visual Encoders
di: Maier, James, et al.
Pubblicazione: (2024)
di: Maier, James, et al.
Pubblicazione: (2024)
TABLET: Table Structure Recognition using Encoder-only Transformers
di: Hou, Qiyu, et al.
Pubblicazione: (2025)
di: Hou, Qiyu, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Zero-Shot Vision Encoder Grafting via LLM Surrogates
di: Yue, Kaiyu, et al.
Pubblicazione: (2025) -
Efficient Image Synthesis with Sphere Latent Encoder
di: Do, Tung, et al.
Pubblicazione: (2026) -
Object Recognition as Next Token Prediction
di: Yue, Kaiyu, et al.
Pubblicazione: (2023) -
From Pixels to Prose: A Large Dataset of Dense Image Captions
di: Singla, Vasu, et al.
Pubblicazione: (2024) -
Language-Image Alignment with Fixed Text Encoders
di: Yang, Jingfeng, et al.
Pubblicazione: (2025)