Dr$^2$Net: Dynamic Reversible Dual-Residual Networks for Memory-Efficient Finetuning
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Zhao, Chen, Liu, Shuming, Mangalam, Karttikeya, Qian, Guocheng, Zohra, Fatimah, Alghannam, Abdulmohsen, Malik, Jitendra, Ghanem, Bernard |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2024
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Adaptive Human Trajectory Prediction via Latent Corridors
von: Thakkar, Neerja, et al.
Veröffentlicht: (2023)
von: Thakkar, Neerja, et al.
Veröffentlicht: (2023)
$β$-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
von: Zohra, Fatimah, et al.
Veröffentlicht: (2025)
von: Zohra, Fatimah, et al.
Veröffentlicht: (2025)
xT: Nested Tokenization for Larger Context in Large Images
von: Gupta, Ritwik, et al.
Veröffentlicht: (2024)
von: Gupta, Ritwik, et al.
Veröffentlicht: (2024)
Pix4Point: Image Pretrained Standard Transformers for 3D Point Cloud Understanding
von: Qian, Guocheng, et al.
Veröffentlicht: (2022)
von: Qian, Guocheng, et al.
Veröffentlicht: (2022)
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
von: Liu, Shuming, et al.
Veröffentlicht: (2025)
von: Liu, Shuming, et al.
Veröffentlicht: (2025)
ColorMAE: Exploring data-independent masking strategies in Masked AutoEncoders
von: Hinojosa, Carlos, et al.
Veröffentlicht: (2024)
von: Hinojosa, Carlos, et al.
Veröffentlicht: (2024)
End-to-End Temporal Action Detection with 1B Parameters Across 1000 Frames
von: Liu, Shuming, et al.
Veröffentlicht: (2023)
von: Liu, Shuming, et al.
Veröffentlicht: (2023)
OpenTAD: A Unified Framework and Comprehensive Study of Temporal Action Detection
von: Liu, Shuming, et al.
Veröffentlicht: (2025)
von: Liu, Shuming, et al.
Veröffentlicht: (2025)
GES: Generalized Exponential Splatting for Efficient Radiance Field Rendering
von: Hamdi, Abdullah, et al.
Veröffentlicht: (2024)
von: Hamdi, Abdullah, et al.
Veröffentlicht: (2024)
Video Self-Stitching Graph Network for Temporal Action Localization
von: Zhao, Chen, et al.
Veröffentlicht: (2020)
von: Zhao, Chen, et al.
Veröffentlicht: (2020)
Mist: Efficient Distributed Training of Large Language Models via Memory-Parallelism Co-Optimization
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
von: Zhu, Zhanda, et al.
Veröffentlicht: (2025)
ResidualViT for Efficient Temporally Dense Video Encoding
von: Soldan, Mattia, et al.
Veröffentlicht: (2025)
von: Soldan, Mattia, et al.
Veröffentlicht: (2025)
Harnessing Temporal Causality for Advanced Temporal Action Detection
von: Liu, Shuming, et al.
Veröffentlicht: (2024)
von: Liu, Shuming, et al.
Veröffentlicht: (2024)
EasyV2V: A High-quality Instruction-based Video Editing Framework
von: Mai, Jinjie, et al.
Veröffentlicht: (2025)
von: Mai, Jinjie, et al.
Veröffentlicht: (2025)
TrackNeRF: Bundle Adjusting NeRF from Sparse and Noisy Views via Feature Tracks
von: Mai, Jinjie, et al.
Veröffentlicht: (2024)
von: Mai, Jinjie, et al.
Veröffentlicht: (2024)
TimeLoc: A Unified End-to-End Framework for Precise Timestamp Localization in Long Videos
von: Zhang, Chen-Lin, et al.
Veröffentlicht: (2025)
von: Zhang, Chen-Lin, et al.
Veröffentlicht: (2025)
DRSI-Net: Dual-Residual Spatial Interaction Network for Multi-Person Pose Estimation
von: Wu, Shang, et al.
Veröffentlicht: (2024)
von: Wu, Shang, et al.
Veröffentlicht: (2024)
GoTTA be Diverse: Rethinking Memory Policies for Test-Time Adaptation
von: Alhuwaider, Shyma, et al.
Veröffentlicht: (2026)
von: Alhuwaider, Shyma, et al.
Veröffentlicht: (2026)
SoccerNet-Caption: Dense Video Captioning for Soccer Broadcasts Commentaries
von: Mkhallati, Hassan, et al.
Veröffentlicht: (2023)
von: Mkhallati, Hassan, et al.
Veröffentlicht: (2023)
Re-evaluating the Need for Multimodal Signals in Unsupervised Grammar Induction
von: Li, Boyi, et al.
Veröffentlicht: (2022)
von: Li, Boyi, et al.
Veröffentlicht: (2022)
Do Vision and Language Encoders Represent the World Similarly?
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2024)
von: Maniparambil, Mayug, et al.
Veröffentlicht: (2024)
DPE-Net: Dual-Parallel Encoder Based Network for Semantic Segmentation of Polyps
von: Manan, Malik Abdul, et al.
Veröffentlicht: (2024)
von: Manan, Malik Abdul, et al.
Veröffentlicht: (2024)
SPAD : Spatially Aware Multiview Diffusers
von: Kant, Yash, et al.
Veröffentlicht: (2024)
von: Kant, Yash, et al.
Veröffentlicht: (2024)
Human-level 3D shape perception emerges from multi-view learning
von: Bonnen, Tyler, et al.
Veröffentlicht: (2026)
von: Bonnen, Tyler, et al.
Veröffentlicht: (2026)
ADVMEM: Adversarial Memory Initialization for Realistic Test-Time Adaptation via Tracklet-Based Benchmarking
von: Alhuwaider, Shyma, et al.
Veröffentlicht: (2025)
von: Alhuwaider, Shyma, et al.
Veröffentlicht: (2025)
Dynamically Masked Discriminator for Generative Adversarial Networks
von: Zhang, Wentian, et al.
Veröffentlicht: (2023)
von: Zhang, Wentian, et al.
Veröffentlicht: (2023)
MambaStyle: Efficient StyleGAN Inversion for Real Image Editing with State-Space Models
von: Lopez, Jhon, et al.
Veröffentlicht: (2025)
von: Lopez, Jhon, et al.
Veröffentlicht: (2025)
SMILE: Infusing Spatial and Motion Semantics in Masked Video Learning
von: Thoker, Fida Mohammad, et al.
Veröffentlicht: (2025)
von: Thoker, Fida Mohammad, et al.
Veröffentlicht: (2025)
DDUNet: Dual Dynamic U-Net for Highly-Efficient Cloud Segmentation
von: Li, Yijie, et al.
Veröffentlicht: (2025)
von: Li, Yijie, et al.
Veröffentlicht: (2025)
Reconstructing Hand-Held Objects in 3D from Images and Videos
von: Wu, Jane, et al.
Veröffentlicht: (2024)
von: Wu, Jane, et al.
Veröffentlicht: (2024)
PraNet-V2: Dual-Supervised Reverse Attention for Medical Image Segmentation
von: Hu, Bo-Cheng, et al.
Veröffentlicht: (2025)
von: Hu, Bo-Cheng, et al.
Veröffentlicht: (2025)
DeGMix: Efficient Multi-Task Dense Prediction with Deformable and Gating Mixer
von: Xu, Yangyang, et al.
Veröffentlicht: (2023)
von: Xu, Yangyang, et al.
Veröffentlicht: (2023)
Hybrid Structure-from-Motion and Camera Relocalization for Enhanced Egocentric Localization
von: Mai, Jinjie, et al.
Veröffentlicht: (2024)
von: Mai, Jinjie, et al.
Veröffentlicht: (2024)
SoccerNet-Tracking: Multiple Object Tracking Dataset and Benchmark in Soccer Videos
von: Cioppa, Anthony, et al.
Veröffentlicht: (2022)
von: Cioppa, Anthony, et al.
Veröffentlicht: (2022)
SPARF: Large-Scale Learning of 3D Sparse Radiance Fields from Few Input Images
von: Hamdi, Abdullah, et al.
Veröffentlicht: (2022)
von: Hamdi, Abdullah, et al.
Veröffentlicht: (2022)
UNet--: Memory-Efficient and Feature-Enhanced Network Architecture based on U-Net with Reduced Skip-Connections
von: Yin, Lingxiao, et al.
Veröffentlicht: (2024)
von: Yin, Lingxiao, et al.
Veröffentlicht: (2024)
Efficient Bayesian Uncertainty Estimation for nnU-Net
von: Zhao, Yidong, et al.
Veröffentlicht: (2022)
von: Zhao, Yidong, et al.
Veröffentlicht: (2022)
SAVeS: Steering Safety Judgments in Vision-Language Models via Semantic Cues
von: Hinojosa, Carlos, et al.
Veröffentlicht: (2026)
von: Hinojosa, Carlos, et al.
Veröffentlicht: (2026)
Mind-the-Glitch: Visual Correspondence for Detecting Inconsistencies in Subject-Driven Generation
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2025)
von: Eldesokey, Abdelrahman, et al.
Veröffentlicht: (2025)
Exploring Missing Modality in Multimodal Egocentric Datasets
von: Ramazanova, Merey, et al.
Veröffentlicht: (2024)
von: Ramazanova, Merey, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
Adaptive Human Trajectory Prediction via Latent Corridors
von: Thakkar, Neerja, et al.
Veröffentlicht: (2023) -
$β$-CLIP: Text-Conditioned Contrastive Learning for Multi-Granular Vision-Language Alignment
von: Zohra, Fatimah, et al.
Veröffentlicht: (2025) -
xT: Nested Tokenization for Larger Context in Large Images
von: Gupta, Ritwik, et al.
Veröffentlicht: (2024) -
Pix4Point: Image Pretrained Standard Transformers for 3D Point Cloud Understanding
von: Qian, Guocheng, et al.
Veröffentlicht: (2022) -
BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
von: Liu, Shuming, et al.
Veröffentlicht: (2025)