TanDiT: Tangent-Plane Diffusion Transformer for High-Quality 360° Panorama Generation
Fuente:
arXiv
Saved in:
| Main Authors: | Çapuk, Hakan, Bond, Andrew, Kızıl, Muhammed Burak, Göçen, Emir, Erdem, Erkut, Erdem, Aykut |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2025)
by: Kizil, Muhammed Burak, et al.
Published: (2025)
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
by: Bond, Andrew, et al.
Published: (2026)
by: Bond, Andrew, et al.
Published: (2026)
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2026)
by: Kizil, Muhammed Burak, et al.
Published: (2026)
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
by: Anees, Abdul Basit, et al.
Published: (2024)
by: Anees, Abdul Basit, et al.
Published: (2024)
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
by: Ercan, Burak, et al.
Published: (2024)
by: Ercan, Burak, et al.
Published: (2024)
EVREAL: Towards a Comprehensive Benchmark and Analysis Suite for Event-based Video Reconstruction
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
GaussianVideo: Efficient Video Representation via Hierarchical Gaussian Splatting
by: Bond, Andrew, et al.
Published: (2025)
by: Bond, Andrew, et al.
Published: (2025)
CLIPAway: Harmonizing Focused Embeddings for Removing Objects via Diffusion Models
by: Ekin, Yigit, et al.
Published: (2024)
by: Ekin, Yigit, et al.
Published: (2024)
HyperE2VID: Improving Event-Based Video Reconstruction via Hypernetworks
by: Ercan, Burak, et al.
Published: (2023)
by: Ercan, Burak, et al.
Published: (2023)
Spherical Vision Transformers for Audio-Visual Saliency Prediction in 360-Degree Videos
by: Cokelek, Mert, et al.
Published: (2025)
by: Cokelek, Mert, et al.
Published: (2025)
SonicDiffusion: Audio-Driven Image Generation and Editing with Pretrained Diffusion Models
by: Biner, Burak Can, et al.
Published: (2024)
by: Biner, Burak Can, et al.
Published: (2024)
A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features
by: Karanfil, Enes, et al.
Published: (2025)
by: Karanfil, Enes, et al.
Published: (2025)
Can Your Model Separate Yolks with a Water Bottle? Benchmarking Physical Commonsense Understanding in Video Generation Models
by: Sanli, Enes, et al.
Published: (2025)
by: Sanli, Enes, et al.
Published: (2025)
VidStyleODE: Disentangled Video Editing via StyleGAN and NeuralODEs
by: Ali, Moayed Haji, et al.
Published: (2023)
by: Ali, Moayed Haji, et al.
Published: (2023)
Evaluating Linguistic Capabilities of Multimodal LLMs in the Lens of Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2024)
by: Dogan, Mustafa, et al.
Published: (2024)
FewMMBench: A Benchmark for Multimodal Few-Shot Learning
by: Dogan, Mustafa, et al.
Published: (2026)
by: Dogan, Mustafa, et al.
Published: (2026)
Object and Relation Centric Representations for Push Effect Prediction
by: Tekden, Ahmet E., et al.
Published: (2021)
by: Tekden, Ahmet E., et al.
Published: (2021)
Sequential Compositional Generalization in Multimodal Models
by: Yagcioglu, Semih, et al.
Published: (2024)
by: Yagcioglu, Semih, et al.
Published: (2024)
Infrared Domain Adaptation with Zero-Shot Quantization
by: Sevsay, Burak, et al.
Published: (2024)
by: Sevsay, Burak, et al.
Published: (2024)
DeVisE: Behavioral Testing of Medical Large Language Models
by: Tagliabue, Camila Zurdo, et al.
Published: (2025)
by: Tagliabue, Camila Zurdo, et al.
Published: (2025)
Learning to Think Like a Cartoon Captionist: Incongruity-Resolution Supervision for Multimodal Humor Understanding
by: Vural, Hatice Merve, et al.
Published: (2026)
by: Vural, Hatice Merve, et al.
Published: (2026)
FuseFormer: A Transformer for Visual and Thermal Image Fusion
by: Erdogan, Aytekin, et al.
Published: (2024)
by: Erdogan, Aytekin, et al.
Published: (2024)
360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion Model
by: Wang, Qian, et al.
Published: (2024)
by: Wang, Qian, et al.
Published: (2024)
Taming Stable Diffusion for Text to 360° Panorama Image Generation
by: Zhang, Cheng, et al.
Published: (2024)
by: Zhang, Cheng, et al.
Published: (2024)
Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer
by: Shao, Ruizhi, et al.
Published: (2024)
by: Shao, Ruizhi, et al.
Published: (2024)
Hippocrates: An Open-Source Framework for Advancing Large Language Models in Healthcare
by: Acikgoz, Emre Can, et al.
Published: (2024)
by: Acikgoz, Emre Can, et al.
Published: (2024)
What Makes for Text to 360-degree Panorama Generation with Stable Diffusion?
by: Ni, Jinhong, et al.
Published: (2025)
by: Ni, Jinhong, et al.
Published: (2025)
PanoDiffusion: 360-degree Panorama Outpainting via Diffusion
by: Wu, Tianhao, et al.
Published: (2023)
by: Wu, Tianhao, et al.
Published: (2023)
SE360: Semantic Edit in 360$^\circ$ Panoramas via Hierarchical Data Construction
by: Zhong, Haoyi, et al.
Published: (2025)
by: Zhong, Haoyi, et al.
Published: (2025)
DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training
by: Feng, Haoran, et al.
Published: (2025)
by: Feng, Haoran, et al.
Published: (2025)
360PanT: Training-Free Text-Driven 360-Degree Panorama-to-Panorama Translation
by: Wang, Hai, et al.
Published: (2024)
by: Wang, Hai, et al.
Published: (2024)
FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less Compute
by: Anagnostidis, Sotiris, et al.
Published: (2025)
by: Anagnostidis, Sotiris, et al.
Published: (2025)
How to Augment for Atmospheric Turbulence Effects on Thermal Adapted Object Detection Models?
by: Uzun, Engin, et al.
Published: (2024)
by: Uzun, Engin, et al.
Published: (2024)
Near-Infrared and Low-Rank Adaptation of Vision Transformers in Remote Sensing
by: Ulku, Irem, et al.
Published: (2024)
by: Ulku, Irem, et al.
Published: (2024)
Enhancing Visual Question Answering through Question-Driven Image Captions as Prompts
by: Özdemir, Övgü, et al.
Published: (2024)
by: Özdemir, Övgü, et al.
Published: (2024)
A Survey on Text-Driven 360-Degree Panorama Generation
by: Wang, Hai, et al.
Published: (2025)
by: Wang, Hai, et al.
Published: (2025)
PixelDiT: Pixel Diffusion Transformers for Image Generation
by: Yu, Yongsheng, et al.
Published: (2025)
by: Yu, Yongsheng, et al.
Published: (2025)
Assessment of the Impact of Fruit Vinegars on the Tenderness and Quality Attributes of Spent Hen Meat
by: Nuran Erdem
Published: (2025)
by: Nuran Erdem
Published: (2025)
From Shock to Strategy: Quality of Life Indicators as Foundations for Post-Disaster Recovery in Türkiye
by: Erdem Ayçiçek
Published: (2025)
by: Erdem Ayçiçek
Published: (2025)
Dense360: Dense Understanding from Omnidirectional Panoramas
by: Zhou, Yikang, et al.
Published: (2025)
by: Zhou, Yikang, et al.
Published: (2025)
Similar Items
-
LAMP: Language-Assisted Motion Planning for Controllable Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2025) -
Beyond Gaussian Bottlenecks: Topologically Aligned Encoding of Vision-Transformer Feature Spaces
by: Bond, Andrew, et al.
Published: (2026) -
Auteur: Language-Driven Cinematographic Framing for Human-Centric Video Generation
by: Kizil, Muhammed Burak, et al.
Published: (2026) -
HyperGAN-CLIP: A Unified Framework for Domain Adaptation, Image Synthesis and Manipulation
by: Anees, Abdul Basit, et al.
Published: (2024) -
HUE Dataset: High-Resolution Event and Frame Sequences for Low-Light Vision
by: Ercan, Burak, et al.
Published: (2024)