Cross Resolution Encoding-Decoding For Detection Transformers
Fuente:
arXiv
Salvato in:
| Autori principali: | Kumar, Ashish, Park, Jaesik |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2024
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Designing Concise ConvNets with Columnar Stages
di: Kumar, Ashish, et al.
Pubblicazione: (2024)
di: Kumar, Ashish, et al.
Pubblicazione: (2024)
High-Speed Stereo Visual SLAM for Low-Powered Computing Devices
di: Kumar, Ashish, et al.
Pubblicazione: (2024)
di: Kumar, Ashish, et al.
Pubblicazione: (2024)
Pick-or-Mix: Dynamic Channel Sampling for ConvNets
di: Kumar, Ashish, et al.
Pubblicazione: (2024)
di: Kumar, Ashish, et al.
Pubblicazione: (2024)
Learning SO(3)-Invariant Semantic Correspondence via Local Shape Transform
di: Park, Chunghyun, et al.
Pubblicazione: (2024)
di: Park, Chunghyun, et al.
Pubblicazione: (2024)
Burst Image Super-Resolution via Multi-Cross Attention Encoding and Multi-Scan State-Space Decoding
di: Huang, Tengda, et al.
Pubblicazione: (2025)
di: Huang, Tengda, et al.
Pubblicazione: (2025)
Exploring Multimodal Diffusion Transformers for Enhanced Prompt-based Image Editing
di: Shin, Joonghyuk, et al.
Pubblicazione: (2025)
di: Shin, Joonghyuk, et al.
Pubblicazione: (2025)
Boosting Resolution Generalization of Diffusion Transformers with Randomized Positional Encodings
di: Hou, Liang, et al.
Pubblicazione: (2025)
di: Hou, Liang, et al.
Pubblicazione: (2025)
InstantDrag: Improving Interactivity in Drag-based Image Editing
di: Shin, Joonghyuk, et al.
Pubblicazione: (2024)
di: Shin, Joonghyuk, et al.
Pubblicazione: (2024)
Leveraging Learned Image Prior for 3D Gaussian Compression
di: Shin, Seungjoo, et al.
Pubblicazione: (2025)
di: Shin, Seungjoo, et al.
Pubblicazione: (2025)
Locality-aware Gaussian Compression for Fast and High-quality Rendering
di: Shin, Seungjoo, et al.
Pubblicazione: (2025)
di: Shin, Seungjoo, et al.
Pubblicazione: (2025)
Distribution Matching Distillation without Fake Score Network
di: Kim, Youngjoong, et al.
Pubblicazione: (2026)
di: Kim, Youngjoong, et al.
Pubblicazione: (2026)
Metropolis-Hastings Sampling for 3D Gaussian Reconstruction
di: Kim, Hyunjin, et al.
Pubblicazione: (2025)
di: Kim, Hyunjin, et al.
Pubblicazione: (2025)
Deep Cost Ray Fusion for Sparse Depth Video Completion
di: Kim, Jungeon, et al.
Pubblicazione: (2024)
di: Kim, Jungeon, et al.
Pubblicazione: (2024)
Targetless LiDAR-Camera Calibration with Neural Gaussian Splatting
di: Jung, Haebeom, et al.
Pubblicazione: (2025)
di: Jung, Haebeom, et al.
Pubblicazione: (2025)
Direct Reward Fine-Tuning on Poses for Single Image to 3D Human in the Wild
di: Do, Seunguk, et al.
Pubblicazione: (2026)
di: Do, Seunguk, et al.
Pubblicazione: (2026)
CodecNeRF: Toward Fast Encoding and Decoding, Compact, and High-quality Novel-view Synthesis
di: Kang, Gyeongjin, et al.
Pubblicazione: (2024)
di: Kang, Gyeongjin, et al.
Pubblicazione: (2024)
Fast Encoding and Decoding for Implicit Video Representation
di: Chen, Hao, et al.
Pubblicazione: (2024)
di: Chen, Hao, et al.
Pubblicazione: (2024)
Extend3D: Town-Scale 3D Generation
di: Yoon, Seungwoo, et al.
Pubblicazione: (2026)
di: Yoon, Seungwoo, et al.
Pubblicazione: (2026)
CF3: Compact and Fast 3D Feature Fields
di: Lee, Hyunjoon, et al.
Pubblicazione: (2025)
di: Lee, Hyunjoon, et al.
Pubblicazione: (2025)
3Doodle: Compact Abstraction of Objects with 3D Strokes
di: Choi, Changwoon, et al.
Pubblicazione: (2024)
di: Choi, Changwoon, et al.
Pubblicazione: (2024)
Recovering Dynamic 3D Sketches from Videos
di: Lee, Jaeah, et al.
Pubblicazione: (2025)
di: Lee, Jaeah, et al.
Pubblicazione: (2025)
360 in the Wild: Dataset for Depth Prediction and View Synthesis
di: Park, Kibaek, et al.
Pubblicazione: (2024)
di: Park, Kibaek, et al.
Pubblicazione: (2024)
Towards Zero-Shot Point Cloud Registration Across Diverse Scales, Scenes, and Sensor Setups
di: Lim, Hyungtae, et al.
Pubblicazione: (2026)
di: Lim, Hyungtae, et al.
Pubblicazione: (2026)
OpenBox: Annotate Any Bounding Boxes in 3D
di: Lee, In-Jae, et al.
Pubblicazione: (2025)
di: Lee, In-Jae, et al.
Pubblicazione: (2025)
Extending CLIP's Image-Text Alignment to Referring Image Segmentation
di: Kim, Seoyeon, et al.
Pubblicazione: (2023)
di: Kim, Seoyeon, et al.
Pubblicazione: (2023)
Coherent Human-Scene Reconstruction from Multi-Person Multi-View Video in a Single Pass
di: Kim, Sangmin, et al.
Pubblicazione: (2026)
di: Kim, Sangmin, et al.
Pubblicazione: (2026)
Cross-Resolution Land Cover Classification Using Outdated Products and Transformers
di: Ni, Huan, et al.
Pubblicazione: (2024)
di: Ni, Huan, et al.
Pubblicazione: (2024)
CFPFormer: Feature-pyramid like Transformer Decoder for Segmentation and Detection
di: Cai, Hongyi, et al.
Pubblicazione: (2024)
di: Cai, Hongyi, et al.
Pubblicazione: (2024)
Improving Editability in Image Generation with Layer-wise Memory
di: Kim, Daneul, et al.
Pubblicazione: (2025)
di: Kim, Daneul, et al.
Pubblicazione: (2025)
Transformer Based Self-Context Aware Prediction for Few-Shot Anomaly Detection in Videos
di: Pillai, Gargi V., et al.
Pubblicazione: (2025)
di: Pillai, Gargi V., et al.
Pubblicazione: (2025)
DBAT: Dynamic Backward Attention Transformer for Material Segmentation with Cross-Resolution Patches
di: Heng, Yuwen, et al.
Pubblicazione: (2023)
di: Heng, Yuwen, et al.
Pubblicazione: (2023)
Characterizing and Optimizing the Spatial Kernel of Multi Resolution Hash Encodings
di: Dai, Tianxiang, et al.
Pubblicazione: (2026)
di: Dai, Tianxiang, et al.
Pubblicazione: (2026)
TERDNet: Transformer Encoder-Recurrent Decoder Network for Scene Change Detection
di: Yoon, Jiae, et al.
Pubblicazione: (2026)
di: Yoon, Jiae, et al.
Pubblicazione: (2026)
CardioCaps: Attention-based Capsule Network for Class-Imbalanced Echocardiogram Classification
di: Han, Hyunkyung, et al.
Pubblicazione: (2024)
di: Han, Hyunkyung, et al.
Pubblicazione: (2024)
Stabilizing Consistency Training: A Flow Map Analysis and Self-Distillation
di: Kim, Youngjoong, et al.
Pubblicazione: (2026)
di: Kim, Youngjoong, et al.
Pubblicazione: (2026)
Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
di: Cao, Guiping, et al.
Pubblicazione: (2025)
di: Cao, Guiping, et al.
Pubblicazione: (2025)
CRA-PCN: Point Cloud Completion with Intra- and Inter-level Cross-Resolution Transformers
di: Rong, Yi, et al.
Pubblicazione: (2024)
di: Rong, Yi, et al.
Pubblicazione: (2024)
Holistic Order Prediction in Natural Scenes
di: Musacchio, Pierre, et al.
Pubblicazione: (2025)
di: Musacchio, Pierre, et al.
Pubblicazione: (2025)
SpatialMosaic: A Multiview VLM Dataset for Partial Visibility
di: Lee, Kanghee, et al.
Pubblicazione: (2025)
di: Lee, Kanghee, et al.
Pubblicazione: (2025)
TRACE: Your Diffusion Model is Secretly an Instance Edge Detector
di: Jo, Sanghyun, et al.
Pubblicazione: (2025)
di: Jo, Sanghyun, et al.
Pubblicazione: (2025)
Documenti analoghi
-
Designing Concise ConvNets with Columnar Stages
di: Kumar, Ashish, et al.
Pubblicazione: (2024) -
High-Speed Stereo Visual SLAM for Low-Powered Computing Devices
di: Kumar, Ashish, et al.
Pubblicazione: (2024) -
Pick-or-Mix: Dynamic Channel Sampling for ConvNets
di: Kumar, Ashish, et al.
Pubblicazione: (2024) -
Learning SO(3)-Invariant Semantic Correspondence via Local Shape Transform
di: Park, Chunghyun, et al.
Pubblicazione: (2024) -
Burst Image Super-Resolution via Multi-Cross Attention Encoding and Multi-Scan State-Space Decoding
di: Huang, Tengda, et al.
Pubblicazione: (2025)