Structured State-Space Regularization for Generation-Friendly Image Tokenization
Fuente:
arXiv
Salvato in:
| Autori principali: | Lee, Jinsung, Oh, Jaemin, Kim, Namhun, Kim, Dongwon, Yoon, Byung-Jun, Kwak, Suha |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2026
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
di: Kim, Dongwon, et al.
Pubblicazione: (2026)
di: Kim, Dongwon, et al.
Pubblicazione: (2026)
Bootstrapping Top-down Information for Self-modulating Slot Attention
di: Kim, Dongwon, et al.
Pubblicazione: (2024)
di: Kim, Dongwon, et al.
Pubblicazione: (2024)
Extending CLIP's Image-Text Alignment to Referring Image Segmentation
di: Kim, Seoyeon, et al.
Pubblicazione: (2023)
di: Kim, Seoyeon, et al.
Pubblicazione: (2023)
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
di: Kim, Dongwon, et al.
Pubblicazione: (2025)
di: Kim, Dongwon, et al.
Pubblicazione: (2025)
PLOT: Text-based Person Search with Part Slot Attention for Corresponding Part Discovery
di: Park, Jicheol, et al.
Pubblicazione: (2024)
di: Park, Jicheol, et al.
Pubblicazione: (2024)
Improving Text-based Person Search via Part-level Cross-modal Correspondence
di: Park, Jicheol, et al.
Pubblicazione: (2024)
di: Park, Jicheol, et al.
Pubblicazione: (2024)
Classification Matters: Improving Video Action Detection with Class-Specific Attention
di: Lee, Jinsung, et al.
Pubblicazione: (2024)
di: Lee, Jinsung, et al.
Pubblicazione: (2024)
Tempered Self-Similarity Alignment for Physically Plausible Video Generation
di: Kim, Manjin, et al.
Pubblicazione: (2026)
di: Kim, Manjin, et al.
Pubblicazione: (2026)
TabFlash: Efficient Table Understanding with Progressive Question Conditioning and Token Focusing
di: Kim, Jongha, et al.
Pubblicazione: (2025)
di: Kim, Jongha, et al.
Pubblicazione: (2025)
FREST: Feature RESToration for Semantic Segmentation under Multiple Adverse Conditions
di: Lee, Sohyun, et al.
Pubblicazione: (2024)
di: Lee, Sohyun, et al.
Pubblicazione: (2024)
GroupCoOp: Group-robust Fine-tuning via Group Prompt Learning
di: Kim, Nayeong, et al.
Pubblicazione: (2025)
di: Kim, Nayeong, et al.
Pubblicazione: (2025)
A-IDE : Agent-Integrated Denoising Experts
di: Cho, Uihyun, et al.
Pubblicazione: (2025)
di: Cho, Uihyun, et al.
Pubblicazione: (2025)
TestDG: Test-time Domain Generalization for Continual Test-time Adaptation
di: Lee, Sohyun, et al.
Pubblicazione: (2025)
di: Lee, Sohyun, et al.
Pubblicazione: (2025)
Part-Aware Bottom-Up Group Reasoning for Fine-Grained Social Interaction Detection
di: Kim, Dongkeun, et al.
Pubblicazione: (2025)
di: Kim, Dongkeun, et al.
Pubblicazione: (2025)
Style-Friendly SNR Sampler for Style-Driven Generation
di: Choi, Jooyoung, et al.
Pubblicazione: (2024)
di: Choi, Jooyoung, et al.
Pubblicazione: (2024)
Learning Unified Distance Metric Across Diverse Data Distributions with Parameter-Efficient Transfer Learning
di: Kim, Sungyeon, et al.
Pubblicazione: (2023)
di: Kim, Sungyeon, et al.
Pubblicazione: (2023)
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
di: Kim, Inho, et al.
Pubblicazione: (2025)
di: Kim, Inho, et al.
Pubblicazione: (2025)
Iterative Prompt Refinement for Safer Text-to-Image Generation
di: Jeon, Jinwoo, et al.
Pubblicazione: (2025)
di: Jeon, Jinwoo, et al.
Pubblicazione: (2025)
Online Temporal Action Localization with Memory-Augmented Transformer
di: Song, Youngkil, et al.
Pubblicazione: (2024)
di: Song, Youngkil, et al.
Pubblicazione: (2024)
Learning Audio-guided Video Representation with Gated Attention for Video-Text Retrieval
di: Jeong, Boseung, et al.
Pubblicazione: (2025)
di: Jeong, Boseung, et al.
Pubblicazione: (2025)
Towards More Practical Group Activity Detection: A New Benchmark and Model
di: Kim, Dongkeun, et al.
Pubblicazione: (2023)
di: Kim, Dongkeun, et al.
Pubblicazione: (2023)
Active Label Correction for Semantic Segmentation with Foundation Models
di: Kim, Hoyoung, et al.
Pubblicazione: (2024)
di: Kim, Hoyoung, et al.
Pubblicazione: (2024)
Extreme Point Supervised Instance Segmentation
di: Lee, Hyeonjun, et al.
Pubblicazione: (2024)
di: Lee, Hyeonjun, et al.
Pubblicazione: (2024)
Efficient and Versatile Robust Fine-Tuning of Zero-shot Models
di: Kim, Sungyeon, et al.
Pubblicazione: (2024)
di: Kim, Sungyeon, et al.
Pubblicazione: (2024)
BioAtt: Anatomical Prior Driven Low-Dose CT Denoising
di: Kim, Namhun, et al.
Pubblicazione: (2025)
di: Kim, Namhun, et al.
Pubblicazione: (2025)
Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding
di: Kim, Jiwan, et al.
Pubblicazione: (2026)
di: Kim, Jiwan, et al.
Pubblicazione: (2026)
GaRA-SAM: Robustifying Segment Anything Model with Gated-Rank Adaptation
di: Lee, Sohyun, et al.
Pubblicazione: (2025)
di: Lee, Sohyun, et al.
Pubblicazione: (2025)
RePL: Pseudo-label Refinement for Semi-supervised LiDAR Semantic Segmentation
di: Kwon, Donghyeon, et al.
Pubblicazione: (2026)
di: Kwon, Donghyeon, et al.
Pubblicazione: (2026)
UCMNet: Uncertainty-Aware Context Memory Network for Under-Display Camera Image Restoration
di: Kim, Daehyun, et al.
Pubblicazione: (2026)
di: Kim, Daehyun, et al.
Pubblicazione: (2026)
Relevance-aware Multi-context Contrastive Decoding for Retrieval-augmented Visual Question Answering
di: Kim, Jongha, et al.
Pubblicazione: (2026)
di: Kim, Jongha, et al.
Pubblicazione: (2026)
Difference Inversion: Interpolate and Isolate the Difference with Token Consistency for Image Analogy Generation
di: Kim, Hyunsoo, et al.
Pubblicazione: (2025)
di: Kim, Hyunsoo, et al.
Pubblicazione: (2025)
Enhancing Multi-Image Understanding through Delimiter Token Scaling
di: Lee, Minyoung, et al.
Pubblicazione: (2026)
di: Lee, Minyoung, et al.
Pubblicazione: (2026)
Enhancing Cost Efficiency in Active Learning with Candidate Set Query
di: Gwon, Yeho, et al.
Pubblicazione: (2025)
di: Gwon, Yeho, et al.
Pubblicazione: (2025)
Improving Robustness to Multiple Spurious Correlations by Multi-Objective Optimization
di: Kim, Nayeong, et al.
Pubblicazione: (2024)
di: Kim, Nayeong, et al.
Pubblicazione: (2024)
Towards Efficient Vision State Space Models via Token Merging
di: Park, Jinyoung, et al.
Pubblicazione: (2025)
di: Park, Jinyoung, et al.
Pubblicazione: (2025)
Generalized Consistency Trajectory Models for Image Manipulation
di: Kim, Beomsu, et al.
Pubblicazione: (2024)
di: Kim, Beomsu, et al.
Pubblicazione: (2024)
Superpixel Tokenization for Vision Transformers: Preserving Semantic Integrity in Visual Tokens
di: Lew, Jaihyun, et al.
Pubblicazione: (2024)
di: Lew, Jaihyun, et al.
Pubblicazione: (2024)
SYNAuG: Exploiting Synthetic Data for Data Imbalance Problems
di: Ye-Bin, Moon, et al.
Pubblicazione: (2023)
di: Ye-Bin, Moon, et al.
Pubblicazione: (2023)
ActFusion: a Unified Diffusion Model for Action Segmentation and Anticipation
di: Gong, Dayoung, et al.
Pubblicazione: (2024)
di: Gong, Dayoung, et al.
Pubblicazione: (2024)
Semantic Prompting with Image-Token for Continual Learning
di: Han, Jisu, et al.
Pubblicazione: (2024)
di: Han, Jisu, et al.
Pubblicazione: (2024)
Documenti analoghi
-
Planning in 8 Tokens: A Compact Discrete Tokenizer for Latent World Model
di: Kim, Dongwon, et al.
Pubblicazione: (2026) -
Bootstrapping Top-down Information for Self-modulating Slot Attention
di: Kim, Dongwon, et al.
Pubblicazione: (2024) -
Extending CLIP's Image-Text Alignment to Referring Image Segmentation
di: Kim, Seoyeon, et al.
Pubblicazione: (2023) -
Democratizing Text-to-Image Masked Generative Models with Compact Text-Aware One-Dimensional Tokens
di: Kim, Dongwon, et al.
Pubblicazione: (2025) -
PLOT: Text-based Person Search with Part Slot Attention for Corresponding Part Discovery
di: Park, Jicheol, et al.
Pubblicazione: (2024)