BIGFix: Bidirectional Image Generation with Token Fixing
Fuente:
arXiv
Saved in:
| Main Authors: | Besnier, Victor, Hurych, David, Bursuc, Andrei, Valle, Eduardo |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Halton Scheduler For Masked Generative Image Transformer
by: Besnier, Victor, et al.
Published: (2025)
by: Besnier, Victor, et al.
Published: (2025)
Supervised Anomaly Detection for Complex Industrial Images
by: Baitieva, Aimira, et al.
Published: (2024)
by: Baitieva, Aimira, et al.
Published: (2024)
POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images
by: Vobecky, Antonin, et al.
Published: (2024)
by: Vobecky, Antonin, et al.
Published: (2024)
Drive&Segment: Unsupervised Semantic Segmentation of Urban Scenes via Cross-modal Distillation
by: Vobecky, Antonin, et al.
Published: (2022)
by: Vobecky, Antonin, et al.
Published: (2022)
VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
by: Bartoccioni, Florent, et al.
Published: (2025)
by: Bartoccioni, Florent, et al.
Published: (2025)
R3DPA: Leveraging 3D Representation Alignment and RGB Pretrained Priors for LiDAR Scene Generation
by: Sereyjol-Garros, Nicolas, et al.
Published: (2026)
by: Sereyjol-Garros, Nicolas, et al.
Published: (2026)
A Simple Recipe for Language-guided Domain Generalized Segmentation
by: Fahes, Mohammad, et al.
Published: (2023)
by: Fahes, Mohammad, et al.
Published: (2023)
Beyond Fixed Topologies: Unregistered Training and Comprehensive Evaluation Metrics for 3D Talking Heads
by: Nocentini, Federico, et al.
Published: (2024)
by: Nocentini, Federico, et al.
Published: (2024)
Don't drop your samples! Coherence-aware training benefits Conditional diffusion
by: Dufour, Nicolas, et al.
Published: (2024)
by: Dufour, Nicolas, et al.
Published: (2024)
Let-It-Flow: Simultaneous Optimization of 3D Flow and Object Clustering
by: Vacek, Patrik, et al.
Published: (2024)
by: Vacek, Patrik, et al.
Published: (2024)
Test-Time Conditioning with Representation-Aligned Visual Features
by: Sereyjol-Garros, Nicolas, et al.
Published: (2026)
by: Sereyjol-Garros, Nicolas, et al.
Published: (2026)
DIP: Unsupervised Dense In-Context Post-training of Visual Representations
by: Sirko-Galouchenko, Sophia, et al.
Published: (2025)
by: Sirko-Galouchenko, Sophia, et al.
Published: (2025)
Boosting Visual Instruction Tuning with Self-Supervised Guidance
by: Sirko-Galouchenko, Sophia, et al.
Published: (2026)
by: Sirko-Galouchenko, Sophia, et al.
Published: (2026)
TokenDance: Token-to-Token Music-to-Dance Generation with Bidirectional Mamba
by: Yang, Ziyue, et al.
Published: (2026)
by: Yang, Ziyue, et al.
Published: (2026)
ScribeTokens: Fixed-Vocabulary Tokenization of Digital Ink
by: Wang, Douglass
Published: (2026)
by: Wang, Douglass
Published: (2026)
LiDAS: Lighting-driven Dynamic Active Sensing for Nighttime Perception
by: de Moreau, Simon, et al.
Published: (2025)
by: de Moreau, Simon, et al.
Published: (2025)
3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
by: Yamada, Ryousuke, et al.
Published: (2025)
by: Yamada, Ryousuke, et al.
Published: (2025)
Improving Semantic Uncertainty Quantification in LVLMs with Semantic Gaussian Processes
by: Hoche, Joseph, et al.
Published: (2025)
by: Hoche, Joseph, et al.
Published: (2025)
CLIP's Visual Embedding Projector is a Few-shot Cornucopia
by: Fahes, Mohammad, et al.
Published: (2024)
by: Fahes, Mohammad, et al.
Published: (2024)
CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation
by: Wysoczańska, Monika, et al.
Published: (2023)
by: Wysoczańska, Monika, et al.
Published: (2023)
Regularizing Self-supervised 3D Scene Flows with Surface Awareness and Cyclic Consistency
by: Vacek, Patrik, et al.
Published: (2023)
by: Vacek, Patrik, et al.
Published: (2023)
ScanMove: Motion Prediction and Transfer for Unregistered Body Meshes
by: Besnier, Thomas, et al.
Published: (2025)
by: Besnier, Thomas, et al.
Published: (2025)
No Train, all Gain: Self-Supervised Gradients Improve Deep Frozen Representations
by: Simoncini, Walter, et al.
Published: (2024)
by: Simoncini, Walter, et al.
Published: (2024)
BiFM: Bidirectional Flow Matching for Few-Step Image Editing and Generation
by: Dai, Yasong, et al.
Published: (2026)
by: Dai, Yasong, et al.
Published: (2026)
Visual-Word Tokenizer: Beyond Fixed Sets of Tokens in Vision Transformers
by: Gee, Leonidas, et al.
Published: (2024)
by: Gee, Leonidas, et al.
Published: (2024)
Learning to Generate Training Datasets for Robust Semantic Segmentation
by: Hariat, Marwane, et al.
Published: (2023)
by: Hariat, Marwane, et al.
Published: (2023)
SuperQuadricOcc: Real-Time Self-Supervised Semantic Occupancy Estimation with Superquadric Volume Rendering
by: Hayes, Seamie, et al.
Published: (2025)
by: Hayes, Seamie, et al.
Published: (2025)
Test-time Contrastive Concepts for Open-world Semantic Segmentation with Vision-Language Models
by: Wysoczańska, Monika, et al.
Published: (2024)
by: Wysoczańska, Monika, et al.
Published: (2024)
OmniTokenizer: A Joint Image-Video Tokenizer for Visual Generation
by: Wang, Junke, et al.
Published: (2024)
by: Wang, Junke, et al.
Published: (2024)
Improving Flexible Image Tokenizers for Autoregressive Image Generation
by: Fu, Zixuan, et al.
Published: (2026)
by: Fu, Zixuan, et al.
Published: (2026)
ImageFolder: Autoregressive Image Generation with Folded Tokens
by: Li, Xiang, et al.
Published: (2024)
by: Li, Xiang, et al.
Published: (2024)
FLOSS: Free Lunch in Open-vocabulary Semantic Segmentation
by: Benigmim, Yasser, et al.
Published: (2025)
by: Benigmim, Yasser, et al.
Published: (2025)
Three Pillars improving Vision Foundation Model Distillation for Lidar
by: Puy, Gilles, et al.
Published: (2023)
by: Puy, Gilles, et al.
Published: (2023)
Domain Adaptation with a Single Vision-Language Embedding
by: Fahes, Mohammad, et al.
Published: (2024)
by: Fahes, Mohammad, et al.
Published: (2024)
Residual-based Efficient Bidirectional Diffusion Model for Image Dehazing and Haze Generation
by: Liu, Bing, et al.
Published: (2025)
by: Liu, Bing, et al.
Published: (2025)
An Image is Worth 32 Tokens for Reconstruction and Generation
by: Yu, Qihang, et al.
Published: (2024)
by: Yu, Qihang, et al.
Published: (2024)
LED: Light Enhanced Depth Estimation at Night
by: de Moreau, Simon, et al.
Published: (2024)
by: de Moreau, Simon, et al.
Published: (2024)
Image Understanding Makes for A Good Tokenizer for Image Generation
by: Wang, Luting, et al.
Published: (2024)
by: Wang, Luting, et al.
Published: (2024)
PaNDaS: Learnable Deformation Modeling with Localized Control
by: Besnier, Thomas, et al.
Published: (2024)
by: Besnier, Thomas, et al.
Published: (2024)
FreeTalk: Emotional Topology-Free 3D Talking Heads
by: Nocentini, Federico, et al.
Published: (2026)
by: Nocentini, Federico, et al.
Published: (2026)
Similar Items
-
Halton Scheduler For Masked Generative Image Transformer
by: Besnier, Victor, et al.
Published: (2025) -
Supervised Anomaly Detection for Complex Industrial Images
by: Baitieva, Aimira, et al.
Published: (2024) -
POP-3D: Open-Vocabulary 3D Occupancy Prediction from Images
by: Vobecky, Antonin, et al.
Published: (2024) -
Drive&Segment: Unsupervised Semantic Segmentation of Urban Scenes via Cross-modal Distillation
by: Vobecky, Antonin, et al.
Published: (2022) -
VaViM and VaVAM: Autonomous Driving through Video Generative Modeling
by: Bartoccioni, Florent, et al.
Published: (2025)