SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Agrawal, Vaibhav, Parihar, Rishubh, Bhat, Pradhaan, Sarvadevabhatla, Ravi Kiran, Babu, R. Venkatesh |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2026
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Compass Control: Multi Object Orientation Control for Text-to-Image Generation
von: Parihar, Rishubh, et al.
Veröffentlicht: (2025)
von: Parihar, Rishubh, et al.
Veröffentlicht: (2025)
MonoPlace3D: Learning 3D-Aware Object Placement for 3D Monocular Detection
von: Parihar, Rishubh, et al.
Veröffentlicht: (2025)
von: Parihar, Rishubh, et al.
Veröffentlicht: (2025)
PreciseControl: Enhancing Text-To-Image Diffusion Models with Fine-Grained Attribute Control
von: Parihar, Rishubh, et al.
Veröffentlicht: (2024)
von: Parihar, Rishubh, et al.
Veröffentlicht: (2024)
Text2Place: Affordance-aware Text Guided Human Placement
von: Parihar, Rishubh, et al.
Veröffentlicht: (2024)
von: Parihar, Rishubh, et al.
Veröffentlicht: (2024)
RoadTones: Tone Controllable Text Generation from Road Event Videos
von: Parikh, Chirag, et al.
Veröffentlicht: (2026)
von: Parikh, Chirag, et al.
Veröffentlicht: (2026)
Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing
von: Parihar, Rishubh, et al.
Veröffentlicht: (2025)
von: Parihar, Rishubh, et al.
Veröffentlicht: (2025)
Balancing Act: Distribution-Guided Debiasing in Diffusion Models
von: Parihar, Rishubh, et al.
Veröffentlicht: (2024)
von: Parihar, Rishubh, et al.
Veröffentlicht: (2024)
Reflecting Reality: Enabling Diffusion Models to Produce Faithful Mirror Reflections
von: Dhiman, Ankit, et al.
Veröffentlicht: (2024)
von: Dhiman, Ankit, et al.
Veröffentlicht: (2024)
OLAF: A Plug-and-Play Framework for Enhanced Multi-object Multi-part Scene Parsing
von: Gupta, Pranav, et al.
Veröffentlicht: (2024)
von: Gupta, Pranav, et al.
Veröffentlicht: (2024)
TexTAR : Textual Attribute Recognition in Multi-domain and Multi-lingual Document Images
von: Kumar, Rohan, et al.
Veröffentlicht: (2025)
von: Kumar, Rohan, et al.
Veröffentlicht: (2025)
MoRAG -- Multi-Fusion Retrieval Augmented Generation for Human Motion
von: Kalakonda, Sai Shashank, et al.
Veröffentlicht: (2024)
von: Kalakonda, Sai Shashank, et al.
Veröffentlicht: (2024)
Unveiling Text in Challenging Stone Inscriptions: A Character-Context-Aware Patching Strategy for Binarization
von: Jena, Pratyush, et al.
Veröffentlicht: (2026)
von: Jena, Pratyush, et al.
Veröffentlicht: (2026)
STRinGS: Selective Text Refinement in Gaussian Splatting
von: Raundhal, Abhinav, et al.
Veröffentlicht: (2025)
von: Raundhal, Abhinav, et al.
Veröffentlicht: (2025)
Transfer-LMR: Heavy-Tail Driving Behavior Recognition in Diverse Traffic Scenarios
von: Parikh, Chirag, et al.
Veröffentlicht: (2024)
von: Parikh, Chirag, et al.
Veröffentlicht: (2024)
DashCop: Automated E-ticket Generation for Two-Wheeler Traffic Violations Using Dashcam Videos
von: Rawat, Deepti, et al.
Veröffentlicht: (2025)
von: Rawat, Deepti, et al.
Veröffentlicht: (2025)
Harnessing Diffusion-Generated Synthetic Images for Fair Image Classification
von: Basu, Abhipsa, et al.
Veröffentlicht: (2025)
von: Basu, Abhipsa, et al.
Veröffentlicht: (2025)
IndicDLP: A Foundational Dataset for Multi-Lingual and Multi-Domain Document Layout Parsing
von: Nath, Oikantik, et al.
Veröffentlicht: (2025)
von: Nath, Oikantik, et al.
Veröffentlicht: (2025)
Controllable Egocentric Video Generation via Occlusion-Aware Sparse 3D Hand Joints
von: Zhang, Chenyangguang, et al.
Veröffentlicht: (2026)
von: Zhang, Chenyangguang, et al.
Veröffentlicht: (2026)
OAHuman: Occlusion-Aware 3D Human Reconstruction from Monocular Images
von: Yang, Yuanwang, et al.
Veröffentlicht: (2026)
von: Yang, Yuanwang, et al.
Veröffentlicht: (2026)
RoadSocial: A Diverse VideoQA Dataset and Benchmark for Road Event Understanding from Social Video Narratives
von: Parikh, Chirag, et al.
Veröffentlicht: (2025)
von: Parikh, Chirag, et al.
Veröffentlicht: (2025)
Seeing Through the Tool: A Controlled Benchmark for Occlusion Robustness in Foundation Segmentation Models
von: Ho, Nhan, et al.
Veröffentlicht: (2026)
von: Ho, Nhan, et al.
Veröffentlicht: (2026)
IDD-X: A Multi-View Dataset for Ego-relative Important Object Localization and Explanation in Dense and Unstructured Traffic
von: Parikh, Chirag, et al.
Veröffentlicht: (2024)
von: Parikh, Chirag, et al.
Veröffentlicht: (2024)
Occlusion-Aware 3D Motion Interpretation for Abnormal Behavior Detection
von: Li, Su, et al.
Veröffentlicht: (2024)
von: Li, Su, et al.
Veröffentlicht: (2024)
Occlusion-aware Text-Image-Point Cloud Pretraining for Open-World 3D Object Recognition
von: Nguyen, Khanh, et al.
Veröffentlicht: (2025)
von: Nguyen, Khanh, et al.
Veröffentlicht: (2025)
Text-Image Conditioned 3D Generation
von: Cen, Jiazhong, et al.
Veröffentlicht: (2026)
von: Cen, Jiazhong, et al.
Veröffentlicht: (2026)
Wave-Former: Through-Occlusion 3D Reconstruction via Wireless Shape Completion
von: Dodds, Laura, et al.
Veröffentlicht: (2025)
von: Dodds, Laura, et al.
Veröffentlicht: (2025)
DreamControl: Control-Based Text-to-3D Generation with 3D Self-Prior
von: Huang, Tianyu, et al.
Veröffentlicht: (2023)
von: Huang, Tianyu, et al.
Veröffentlicht: (2023)
DehazeGS: Seeing Through Fog with 3D Gaussian Splatting
von: Yu, Jinze, et al.
Veröffentlicht: (2025)
von: Yu, Jinze, et al.
Veröffentlicht: (2025)
Lookalike3D: Seeing Double in 3D
von: Yeshwanth, Chandan, et al.
Veröffentlicht: (2026)
von: Yeshwanth, Chandan, et al.
Veröffentlicht: (2026)
Occlusion-Aware 3D Hand-Object Pose Estimation with Masked AutoEncoders
von: Yang, Hui, et al.
Veröffentlicht: (2025)
von: Yang, Hui, et al.
Veröffentlicht: (2025)
Shoot-Bounce-3D: Single-Shot Occlusion-Aware 3D from Lidar by Decomposing Two-Bounce Light
von: Klinghoffer, Tzofi, et al.
Veröffentlicht: (2025)
von: Klinghoffer, Tzofi, et al.
Veröffentlicht: (2025)
PI3D: Efficient Text-to-3D Generation with Pseudo-Image Diffusion
von: Liu, Ying-Tian, et al.
Veröffentlicht: (2023)
von: Liu, Ying-Tian, et al.
Veröffentlicht: (2023)
CineMaster: A 3D-Aware and Controllable Framework for Cinematic Text-to-Video Generation
von: Wang, Qinghe, et al.
Veröffentlicht: (2025)
von: Wang, Qinghe, et al.
Veröffentlicht: (2025)
CrackUDA: Incremental Unsupervised Domain Adaptation for Improved Crack Segmentation in Civil Structures
von: Srivastava, Kushagra, et al.
Veröffentlicht: (2024)
von: Srivastava, Kushagra, et al.
Veröffentlicht: (2024)
DreamCS: Geometry-Aware Text-to-3D Generation with Unpaired 3D Reward Supervision
von: Zou, Xiandong, et al.
Veröffentlicht: (2025)
von: Zou, Xiandong, et al.
Veröffentlicht: (2025)
UniC-Lift: Unified 3D Instance Segmentation via Contrastive Learning
von: Dhiman, Ankit, et al.
Veröffentlicht: (2025)
von: Dhiman, Ankit, et al.
Veröffentlicht: (2025)
SplatFont3D: Structure-Aware Text-to-3D Artistic Font Generation with Part-Level Style Control
von: Gan, Ji, et al.
Veröffentlicht: (2025)
von: Gan, Ji, et al.
Veröffentlicht: (2025)
Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
von: Go, Hyojun, et al.
Veröffentlicht: (2025)
von: Go, Hyojun, et al.
Veröffentlicht: (2025)
Generalizing Single-View 3D Shape Retrieval to Occlusions and Unseen Objects
von: Wu, Qirui, et al.
Veröffentlicht: (2023)
von: Wu, Qirui, et al.
Veröffentlicht: (2023)
SIC3D: Style Image Conditioned Text-to-3D Gaussian Splatting Generation
von: He, Ming, et al.
Veröffentlicht: (2026)
von: He, Ming, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
Compass Control: Multi Object Orientation Control for Text-to-Image Generation
von: Parihar, Rishubh, et al.
Veröffentlicht: (2025) -
MonoPlace3D: Learning 3D-Aware Object Placement for 3D Monocular Detection
von: Parihar, Rishubh, et al.
Veröffentlicht: (2025) -
PreciseControl: Enhancing Text-To-Image Diffusion Models with Fine-Grained Attribute Control
von: Parihar, Rishubh, et al.
Veröffentlicht: (2024) -
Text2Place: Affordance-aware Text Guided Human Placement
von: Parihar, Rishubh, et al.
Veröffentlicht: (2024) -
RoadTones: Tone Controllable Text Generation from Road Event Videos
von: Parikh, Chirag, et al.
Veröffentlicht: (2026)