GIVT: Generative Infinite-Vocabulary Transformers
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Tschannen, Michael, Eastwood, Cian, Mentzer, Fabian |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Elucidating the Design Space of Flow Matching for Cellular Microscopy
von: Jones, Charles, et al.
Veröffentlicht: (2026)
von: Jones, Charles, et al.
Veröffentlicht: (2026)
Jet: A Modern Transformer-Based Normalizing Flow
von: Kolesnikov, Alexander, et al.
Veröffentlicht: (2024)
von: Kolesnikov, Alexander, et al.
Veröffentlicht: (2024)
JetFormer: An Autoregressive Generative Model of Raw Images and Text
von: Tschannen, Michael, et al.
Veröffentlicht: (2024)
von: Tschannen, Michael, et al.
Veröffentlicht: (2024)
Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers
von: Stanić, Aleksandar, et al.
Veröffentlicht: (2024)
von: Stanić, Aleksandar, et al.
Veröffentlicht: (2024)
MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice
von: Yi, Hongwei, et al.
Veröffentlicht: (2025)
von: Yi, Hongwei, et al.
Veröffentlicht: (2025)
High-Fidelity Image Compression with Score-based Generative Models
von: Hoogeboom, Emiel, et al.
Veröffentlicht: (2023)
von: Hoogeboom, Emiel, et al.
Veröffentlicht: (2023)
Scaling Training Data with Lossy Image Compression
von: Mentzer, Katherine L., et al.
Veröffentlicht: (2024)
von: Mentzer, Katherine L., et al.
Veröffentlicht: (2024)
DiffTF++: 3D-aware Diffusion Transformer for Large-Vocabulary 3D Generation
von: Cao, Ziang, et al.
Veröffentlicht: (2024)
von: Cao, Ziang, et al.
Veröffentlicht: (2024)
LangOcc: Self-Supervised Open Vocabulary Occupancy Estimation via Volume Rendering
von: Boeder, Simon, et al.
Veröffentlicht: (2024)
von: Boeder, Simon, et al.
Veröffentlicht: (2024)
Semantic Segmentation and Scene Reconstruction of RGB-D Image Frames: An End-to-End Modular Pipeline for Robotic Applications
von: Zheng, Zhiwu, et al.
Veröffentlicht: (2024)
von: Zheng, Zhiwu, et al.
Veröffentlicht: (2024)
OVLW-DETR: Open-Vocabulary Light-Weighted Detection Transformer
von: Wang, Yu, et al.
Veröffentlicht: (2024)
von: Wang, Yu, et al.
Veröffentlicht: (2024)
InfiniteVGGT: Visual Geometry Grounded Transformer for Endless Streams
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)
von: Yuan, Shuai, et al.
Veröffentlicht: (2026)
CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction
von: Wu, Size, et al.
Veröffentlicht: (2023)
von: Wu, Size, et al.
Veröffentlicht: (2023)
Generalization Boosted Adapter for Open-Vocabulary Segmentation
von: Xu, Wenhao, et al.
Veröffentlicht: (2024)
von: Xu, Wenhao, et al.
Veröffentlicht: (2024)
OVTR: End-to-End Open-Vocabulary Multiple Object Tracking with Transformer
von: Li, Jinyang, et al.
Veröffentlicht: (2025)
von: Li, Jinyang, et al.
Veröffentlicht: (2025)
Infinite Gaze Generation for Videos with Autoregressive Diffusion
von: Kang, Jenna, et al.
Veröffentlicht: (2026)
von: Kang, Jenna, et al.
Veröffentlicht: (2026)
Open-Vocabulary Domain Generalization in Urban-Scene Segmentation
von: Zhao, Dong, et al.
Veröffentlicht: (2026)
von: Zhao, Dong, et al.
Veröffentlicht: (2026)
From Open-Vocabulary to Vocabulary-Free Semantic Segmentation
von: Reichard, Klara, et al.
Veröffentlicht: (2025)
von: Reichard, Klara, et al.
Veröffentlicht: (2025)
Self-Supervised Disentanglement by Leveraging Structure in Data Augmentations
von: Eastwood, Cian, et al.
Veröffentlicht: (2023)
von: Eastwood, Cian, et al.
Veröffentlicht: (2023)
Seg4Diff: Unveiling Open-Vocabulary Segmentation in Text-to-Image Diffusion Transformers
von: Kim, Chaehyun, et al.
Veröffentlicht: (2025)
von: Kim, Chaehyun, et al.
Veröffentlicht: (2025)
CityGen: Infinite and Controllable City Layout Generation
von: Deng, Jie, et al.
Veröffentlicht: (2023)
von: Deng, Jie, et al.
Veröffentlicht: (2023)
Learning to Generalize without Bias for Open-Vocabulary Action Recognition
von: Yu, Yating, et al.
Veröffentlicht: (2025)
von: Yu, Yating, et al.
Veröffentlicht: (2025)
VLog: Video-Language Models by Generative Retrieval of Narration Vocabulary
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
von: Lin, Kevin Qinghong, et al.
Veröffentlicht: (2025)
Self-Prompting Diffusion Transformer for Open-Vocabulary Scene Text Editing via In-Context Learning
von: Li, Hongxi, et al.
Veröffentlicht: (2026)
von: Li, Hongxi, et al.
Veröffentlicht: (2026)
LocCa: Visual Pretraining with Location-aware Captioners
von: Wan, Bo, et al.
Veröffentlicht: (2024)
von: Wan, Bo, et al.
Veröffentlicht: (2024)
Infinite-Homography as Robust Conditioning for Camera-Controlled Video Generation
von: Kim, Min-Jung, et al.
Veröffentlicht: (2025)
von: Kim, Min-Jung, et al.
Veröffentlicht: (2025)
RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection
von: Chen, Fangyi, et al.
Veröffentlicht: (2024)
von: Chen, Fangyi, et al.
Veröffentlicht: (2024)
Leveraging Depth and Language for Open-Vocabulary Domain-Generalized Semantic Segmentation
von: Chen, Siyu, et al.
Veröffentlicht: (2025)
von: Chen, Siyu, et al.
Veröffentlicht: (2025)
Towards scientific discovery with dictionary learning: Extracting biological concepts from microscopy foundation models
von: Donhauser, Konstantin, et al.
Veröffentlicht: (2024)
von: Donhauser, Konstantin, et al.
Veröffentlicht: (2024)
Synthetic Captions for Open-Vocabulary Zero-Shot Segmentation
von: Lebailly, Tim, et al.
Veröffentlicht: (2025)
von: Lebailly, Tim, et al.
Veröffentlicht: (2025)
Real-time Transformer-based Open-Vocabulary Detection with Efficient Fusion Head
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2024)
von: Zhao, Tiancheng, et al.
Veröffentlicht: (2024)
Lost in Translation? Vocabulary Alignment for Source-Free Adaptation in Open-Vocabulary Semantic Segmentation
von: Mazzucco, Silvio, et al.
Veröffentlicht: (2025)
von: Mazzucco, Silvio, et al.
Veröffentlicht: (2025)
SkyReels-V2: Infinite-length Film Generative Model
von: Chen, Guibin, et al.
Veröffentlicht: (2025)
von: Chen, Guibin, et al.
Veröffentlicht: (2025)
Infinite Motion: Extended Motion Generation via Long Text Instructions
von: Li, Mengtian, et al.
Veröffentlicht: (2024)
von: Li, Mengtian, et al.
Veröffentlicht: (2024)
StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025)
von: Tu, Shuyuan, et al.
Veröffentlicht: (2025)
Infinite-Story: A Training-Free Consistent Text-to-Image Generation
von: Park, Jihun, et al.
Veröffentlicht: (2025)
von: Park, Jihun, et al.
Veröffentlicht: (2025)
Stable Video Infinity: Infinite-Length Video Generation with Error Recycling
von: Li, Wuyang, et al.
Veröffentlicht: (2025)
von: Li, Wuyang, et al.
Veröffentlicht: (2025)
Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos
von: Feng, X., et al.
Veröffentlicht: (2026)
von: Feng, X., et al.
Veröffentlicht: (2026)
Affogato: Learning Open-Vocabulary Affordance Grounding with Automated Data Generation at Scale
von: Lee, Junha, et al.
Veröffentlicht: (2025)
von: Lee, Junha, et al.
Veröffentlicht: (2025)
GHOST: Grounded Human Motion Generation with Open Vocabulary Scene-and-Text Contexts
von: Milacski, Zoltán Á., et al.
Veröffentlicht: (2024)
von: Milacski, Zoltán Á., et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Elucidating the Design Space of Flow Matching for Cellular Microscopy
von: Jones, Charles, et al.
Veröffentlicht: (2026) -
Jet: A Modern Transformer-Based Normalizing Flow
von: Kolesnikov, Alexander, et al.
Veröffentlicht: (2024) -
JetFormer: An Autoregressive Generative Model of Raw Images and Text
von: Tschannen, Michael, et al.
Veröffentlicht: (2024) -
Towards Truly Zero-shot Compositional Visual Reasoning with LLMs as Programmers
von: Stanić, Aleksandar, et al.
Veröffentlicht: (2024) -
MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice
von: Yi, Hongwei, et al.
Veröffentlicht: (2025)