Saved in:
| Main Authors: | Yang, Yuchen, Wang, Xinyi, Li, Dong, Tian, Lu, Sirasao, Ashish, Yang, Xun |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2407.10406 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
UPDP: A Unified Progressive Depth Pruner for CNN and Vision Transformer
by: Liu, Ji, et al.
Published: (2024)
by: Liu, Ji, et al.
Published: (2024)
CLAIM: Camera-LiDAR Alignment with Intensity and Monodepth
by: Zhang, Zhuo, et al.
Published: (2025)
by: Zhang, Zhuo, et al.
Published: (2025)
Sparse Laneformer
by: Liu, Ji, et al.
Published: (2024)
by: Liu, Ji, et al.
Published: (2024)
MonSter++: Unified Stereo Matching, Multi-view Stereo, and Real-time Stereo with Monodepth Priors
by: Cheng, Junda, et al.
Published: (2025)
by: Cheng, Junda, et al.
Published: (2025)
AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models
by: Wang, Xinyi, et al.
Published: (2025)
by: Wang, Xinyi, et al.
Published: (2025)
Adjacent-view Transformers for Supervised Surround-view Depth Estimation
by: Guo, Xianda, et al.
Published: (2023)
by: Guo, Xianda, et al.
Published: (2023)
Scale-invariant and View-relational Representation Learning for Full Surround Monocular Depth
by: Hwang, Kyumin, et al.
Published: (2025)
by: Hwang, Kyumin, et al.
Published: (2025)
End-to-End Multi-Person Pose Estimation with Pose-Aware Video Transformer
by: Yu, Yonghui, et al.
Published: (2025)
by: Yu, Yonghui, et al.
Published: (2025)
DiP-GO: A Diffusion Pruner via Few-step Gradient Optimization
by: Zhu, Haowei, et al.
Published: (2024)
by: Zhu, Haowei, et al.
Published: (2024)
SurroundNet: Towards Effective Low-Light Image Enhancement
by: Zhou, Fei, et al.
Published: (2021)
by: Zhou, Fei, et al.
Published: (2021)
SurroundSDF: Implicit 3D Scene Understanding Based on Signed Distance Field
by: Liu, Lizhe, et al.
Published: (2024)
by: Liu, Lizhe, et al.
Published: (2024)
Towards Scale-Aware Low-Light Enhancement via Structure-Guided Transformer Design
by: Dong, Wei, et al.
Published: (2025)
by: Dong, Wei, et al.
Published: (2025)
DrivingGaussian++: Towards Realistic Reconstruction and Editable Simulation for Surrounding Dynamic Driving Scenes
by: Xiong, Yajiao, et al.
Published: (2025)
by: Xiong, Yajiao, et al.
Published: (2025)
Amber-Image: Efficient Compression of Large-Scale Diffusion Transformers
by: Yang, Chaojie, et al.
Published: (2026)
by: Yang, Chaojie, et al.
Published: (2026)
FitControler: Toward Fit-Aware Virtual Try-On
by: Yang, Lu, et al.
Published: (2025)
by: Yang, Lu, et al.
Published: (2025)
Towards Cross-View-Consistent Self-Supervised Surround Depth Estimation
by: Ding, Laiyan, et al.
Published: (2024)
by: Ding, Laiyan, et al.
Published: (2024)
AugRefer: Advancing 3D Visual Grounding via Cross-Modal Augmentation and Spatial Relation-based Referring
by: Wang, Xinyi, et al.
Published: (2025)
by: Wang, Xinyi, et al.
Published: (2025)
Future-Aware Interaction Network For Motion Forecasting
by: Li, Shijie, et al.
Published: (2025)
by: Li, Shijie, et al.
Published: (2025)
Towards Fine-grained Renal Vasculature Segmentation: Full-Scale Hierarchical Learning with FH-Seg
by: Long, Yitian, et al.
Published: (2025)
by: Long, Yitian, et al.
Published: (2025)
FishBEV: Distortion-Resilient Bird's Eye View Segmentation with Surround-View Fisheye Cameras
by: Li, Hang, et al.
Published: (2025)
by: Li, Hang, et al.
Published: (2025)
Transformer Based Self-Context Aware Prediction for Few-Shot Anomaly Detection in Videos
by: Pillai, Gargi V., et al.
Published: (2025)
by: Pillai, Gargi V., et al.
Published: (2025)
Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation
by: Fang, Hongwei, et al.
Published: (2026)
by: Fang, Hongwei, et al.
Published: (2026)
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering
by: Zhou, Sheng, et al.
Published: (2025)
by: Zhou, Sheng, et al.
Published: (2025)
FullTransNet: Full Transformer with Local-Global Attention for Video Summarization
by: Lan, Libin, et al.
Published: (2025)
by: Lan, Libin, et al.
Published: (2025)
DrivingGaussian: Composite Gaussian Splatting for Surrounding Dynamic Autonomous Driving Scenes
by: Zhou, Xiaoyu, et al.
Published: (2023)
by: Zhou, Xiaoyu, et al.
Published: (2023)
FEAT: Full-Dimensional Efficient Attention Transformer for Medical Video Generation
by: Wang, Huihan, et al.
Published: (2025)
by: Wang, Huihan, et al.
Published: (2025)
DEGAS: Detailed Expressions on Full-Body Gaussian Avatars
by: Shao, Zhijing, et al.
Published: (2024)
by: Shao, Zhijing, et al.
Published: (2024)
PAINT: Pathology-Aware Integrated Next-Scale Transformation for Virtual Immunohistochemistry
by: Ma, Rongze, et al.
Published: (2026)
by: Ma, Rongze, et al.
Published: (2026)
GATS: Gaussian Aware Temporal Scaling Transformer for Invariant 4D Spatio-Temporal Point Cloud Representation
by: Tian, Jiayi, et al.
Published: (2026)
by: Tian, Jiayi, et al.
Published: (2026)
In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
by: Zhang, Zechuan, et al.
Published: (2025)
by: Zhang, Zechuan, et al.
Published: (2025)
CronusVLA: Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
by: Li, Hao, et al.
Published: (2025)
by: Li, Hao, et al.
Published: (2025)
SurANet: Surrounding-Aware Network for Concealed Object Detection via Highly-Efficient Interactive Contrastive Learning Strategy
by: Kang, Yuhan, et al.
Published: (2024)
by: Kang, Yuhan, et al.
Published: (2024)
A Timely Survey on Vision Transformer for Deepfake Detection
by: Wang, Zhikan, et al.
Published: (2024)
by: Wang, Zhikan, et al.
Published: (2024)
UniPR-3D: Towards Universal Visual Place Recognition with Visual Geometry Grounded Transformer
by: Deng, Tianchen, et al.
Published: (2025)
by: Deng, Tianchen, et al.
Published: (2025)
PRVR: Partially Relevant Video Retrieval
by: Chen, Xianke, et al.
Published: (2022)
by: Chen, Xianke, et al.
Published: (2022)
Q-DiT4SR: Exploration of Detail-Preserving Diffusion Transformer Quantization for Real-World Image Super-Resolution
by: Zhang, Xun, et al.
Published: (2026)
by: Zhang, Xun, et al.
Published: (2026)
SeedVR: Seeding Infinity in Diffusion Transformer Towards Generic Video Restoration
by: Wang, Jianyi, et al.
Published: (2025)
by: Wang, Jianyi, et al.
Published: (2025)
DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers
by: Yang, Mengping, et al.
Published: (2026)
by: Yang, Mengping, et al.
Published: (2026)
Global-Aware Monocular Semantic Scene Completion with State Space Models
by: Li, Shijie, et al.
Published: (2025)
by: Li, Shijie, et al.
Published: (2025)
Motion-Aware Transformer for Multi-Object Tracking
by: Yang, Xu, et al.
Published: (2025)
by: Yang, Xu, et al.
Published: (2025)
Similar Items
-
UPDP: A Unified Progressive Depth Pruner for CNN and Vision Transformer
by: Liu, Ji, et al.
Published: (2024) -
CLAIM: Camera-LiDAR Alignment with Intensity and Monodepth
by: Zhang, Zhuo, et al.
Published: (2025) -
Sparse Laneformer
by: Liu, Ji, et al.
Published: (2024) -
MonSter++: Unified Stereo Matching, Multi-view Stereo, and Real-time Stereo with Monodepth Priors
by: Cheng, Junda, et al.
Published: (2025) -
AffordBot: 3D Fine-grained Embodied Reasoning via Multimodal Large Language Models
by: Wang, Xinyi, et al.
Published: (2025)