Flowing from Words to Pixels: A Noise-Free Framework for Cross-Modality Evolution
Fuente:
arXiv
Guardado en:
| Autores principales: | Liu, Qihao, Yin, Xi, Yuille, Alan, Brown, Andrew, Singh, Mannat |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D Data
por: Liu, Qihao, et al.
Publicado: (2024)
por: Liu, Qihao, et al.
Publicado: (2024)
From Pixels to Objects: A Hierarchical Approach for Part and Object Segmentation Using Local and Global Aggregation
por: Xie, Yunfei, et al.
Publicado: (2024)
por: Xie, Yunfei, et al.
Publicado: (2024)
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
por: Liu, Qihao, et al.
Publicado: (2025)
por: Liu, Qihao, et al.
Publicado: (2025)
From Pixels and Words to Waves: A Unified Framework for Spectral Dictionary vLLMs
por: Kiruluta, Andrew, et al.
Publicado: (2025)
por: Kiruluta, Andrew, et al.
Publicado: (2025)
Computer Vision and Its Relationship to Cognitive Science: A perspective from Bayes Decision Theory
por: Yuille, Alan, et al.
Publicado: (2026)
por: Yuille, Alan, et al.
Publicado: (2026)
Differences That Matter: Auditing Models for Capability Gap Discovery and Rectification
por: Liu, Qihao, et al.
Publicado: (2025)
por: Liu, Qihao, et al.
Publicado: (2025)
Gaussian Scenes: Pose-Free Sparse-View Scene Reconstruction using Depth-Enhanced Diffusion Priors
por: Paul, Soumava, et al.
Publicado: (2024)
por: Paul, Soumava, et al.
Publicado: (2024)
TriDiff-4D: Fast 4D Generation through Diffusion-based Triplane Re-posing
por: Sheung, Eddie Pokming, et al.
Publicado: (2025)
por: Sheung, Eddie Pokming, et al.
Publicado: (2025)
Rethinking Video-Text Understanding: Retrieval from Counterfactually Augmented Data
por: Ma, Wufei, et al.
Publicado: (2024)
por: Ma, Wufei, et al.
Publicado: (2024)
From Pixel to Cancer: Cellular Automata in Computed Tomography
por: Lai, Yuxiang, et al.
Publicado: (2024)
por: Lai, Yuxiang, et al.
Publicado: (2024)
ImageNet3D: Towards General-Purpose Object-Level 3D Understanding
por: Ma, Wufei, et al.
Publicado: (2024)
por: Ma, Wufei, et al.
Publicado: (2024)
ViMix-14M: A Curated Multi-Source Video-Text Dataset with Long-Form, High-Quality Captions and Crawl-Free Access
por: Yang, Timing, et al.
Publicado: (2025)
por: Yang, Timing, et al.
Publicado: (2025)
Dictionary-based Framework for Interpretable and Consistent Object Parsing
por: Zhang, Tiezheng, et al.
Publicado: (2025)
por: Zhang, Tiezheng, et al.
Publicado: (2025)
SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning
por: Ma, Wufei, et al.
Publicado: (2025)
por: Ma, Wufei, et al.
Publicado: (2025)
4D-Animal: Freely Reconstructing Animatable 3D Animals from Videos
por: Zhong, Shanshan, et al.
Publicado: (2025)
por: Zhong, Shanshan, et al.
Publicado: (2025)
CamFreeDiff: Camera-free Image to Panorama Generation with Diffusion Model
por: Yuan, Xiaoding, et al.
Publicado: (2024)
por: Yuan, Xiaoding, et al.
Publicado: (2024)
XModBench: Benchmarking Cross-Modal Capabilities and Consistency in Omni-Language Models
por: Wang, Xingrui, et al.
Publicado: (2025)
por: Wang, Xingrui, et al.
Publicado: (2025)
FlowAR: Scale-wise Autoregressive Image Generation Meets Flow Matching
por: Ren, Sucheng, et al.
Publicado: (2024)
por: Ren, Sucheng, et al.
Publicado: (2024)
PixelFlow: Pixel-Space Generative Models with Flow
por: Chen, Shoufa, et al.
Publicado: (2025)
por: Chen, Shoufa, et al.
Publicado: (2025)
Grouping First, Attending Smartly: Training-Free Acceleration for Diffusion Transformers
por: Ren, Sucheng, et al.
Publicado: (2025)
por: Ren, Sucheng, et al.
Publicado: (2025)
Generative Animations: A Multi-Model Pipeline for Prompt-Driven Motion Synthesis
por: Khurana, Mannat, et al.
Publicado: (2026)
por: Khurana, Mannat, et al.
Publicado: (2026)
Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate
por: Paul, Soumava, et al.
Publicado: (2026)
por: Paul, Soumava, et al.
Publicado: (2026)
SCLIP: Rethinking Self-Attention for Dense Vision-Language Inference
por: Wang, Feng, et al.
Publicado: (2023)
por: Wang, Feng, et al.
Publicado: (2023)
Quality Sentinel: Estimating Label Quality and Errors in Medical Segmentation Datasets
por: Chen, Yixiong, et al.
Publicado: (2024)
por: Chen, Yixiong, et al.
Publicado: (2024)
Source-Free and Image-Only Unsupervised Domain Adaptation for Category Level Object Pose Estimation
por: Kaushik, Prakhar, et al.
Publicado: (2024)
por: Kaushik, Prakhar, et al.
Publicado: (2024)
EchoingPixels: Cross-Modal Adaptive Token Reduction for Efficient Audio-Visual LLMs
por: Gong, Chao, et al.
Publicado: (2025)
por: Gong, Chao, et al.
Publicado: (2025)
FlowTok: Flowing Seamlessly Across Text and Image Tokens
por: He, Ju, et al.
Publicado: (2025)
por: He, Ju, et al.
Publicado: (2025)
LychSim: A Controllable and Interactive Simulation Framework for Vision Research
por: Ma, Wufei, et al.
Publicado: (2026)
por: Ma, Wufei, et al.
Publicado: (2026)
Frequency-Aware Flow Matching for High-Quality Image Generation
por: Ren, Sucheng, et al.
Publicado: (2026)
por: Ren, Sucheng, et al.
Publicado: (2026)
Beyond Masks: The Case for Medical Image Parsing
por: Gupta, Siddharth, et al.
Publicado: (2026)
por: Gupta, Siddharth, et al.
Publicado: (2026)
RNN as Linear Transformer: A Closer Investigation into Representational Potentials of Visual Mamba Models
por: Yang, Timing, et al.
Publicado: (2025)
por: Yang, Timing, et al.
Publicado: (2025)
From Pixels to Words -- Towards Native One-Vision Models at Scale
por: Diao, Haiwen, et al.
Publicado: (2026)
por: Diao, Haiwen, et al.
Publicado: (2026)
CrossFlowDG: Bridging the Modality Gap with Cross-modal Flow Matching for Domain Generalization
por: Kritikos, Antonios, et al.
Publicado: (2026)
por: Kritikos, Antonios, et al.
Publicado: (2026)
Cross-Modal Prototype Alignment and Mixing for Training-Free Few-Shot Classification
por: Goswami, Dipam, et al.
Publicado: (2026)
por: Goswami, Dipam, et al.
Publicado: (2026)
A Bayesian Approach to OOD Robustness in Image Classification
por: Kaushik, Prakhar, et al.
Publicado: (2024)
por: Kaushik, Prakhar, et al.
Publicado: (2024)
Exploring Cross-Modal Flows for Few-Shot Learning
por: Jiang, Ziqi, et al.
Publicado: (2025)
por: Jiang, Ziqi, et al.
Publicado: (2025)
PixelFlowCast: Latent-Free Precipitation Nowcasting via Pixel Mean Flows
por: Zhu, Yufeng, et al.
Publicado: (2026)
por: Zhu, Yufeng, et al.
Publicado: (2026)
Cross-Modal Mapping: Mitigating the Modality Gap for Few-Shot Image Classification
por: Yang, Xi, et al.
Publicado: (2024)
por: Yang, Xi, et al.
Publicado: (2024)
Cross-Modal Knowledge Distillation for PET-Free Amyloid-Beta Detection from MRI
por: Chiumento, Francesco, et al.
Publicado: (2026)
por: Chiumento, Francesco, et al.
Publicado: (2026)
Fast-then-Fine: A Two-Stage Framework with Multi-Granular Representation for Cross-Modal Retrieval in Remote Sensing
por: Chen, Xi, et al.
Publicado: (2026)
por: Chen, Xi, et al.
Publicado: (2026)
Ejemplares similares
-
DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D Data
por: Liu, Qihao, et al.
Publicado: (2024) -
From Pixels to Objects: A Hierarchical Approach for Part and Object Segmentation Using Local and Global Aggregation
por: Xie, Yunfei, et al.
Publicado: (2024) -
ReVision: Refining Video Diffusion with Explicit 3D Motion Modeling
por: Liu, Qihao, et al.
Publicado: (2025) -
From Pixels and Words to Waves: A Unified Framework for Spectral Dictionary vLLMs
por: Kiruluta, Andrew, et al.
Publicado: (2025) -
Computer Vision and Its Relationship to Cognitive Science: A perspective from Bayes Decision Theory
por: Yuille, Alan, et al.
Publicado: (2026)