Saved in:
| Main Authors: | Tian, Boyuan, Pang, Yihan, Huzaifa, Muhammad, Wang, Shenlong, Adve, Sarita |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2409.04018 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
by: Huzaifa, Muhammad, et al.
Published: (2024)
by: Huzaifa, Muhammad, et al.
Published: (2024)
CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets
by: Amangeldi, Aidar, et al.
Published: (2025)
by: Amangeldi, Aidar, et al.
Published: (2025)
PhysGen3D: Crafting a Miniature Interactive World from a Single Image
by: Chen, Boyuan, et al.
Published: (2025)
by: Chen, Boyuan, et al.
Published: (2025)
Towards Latency-Aware 3D Streaming Perception for Autonomous Driving
by: Peng, Jiaqi, et al.
Published: (2025)
by: Peng, Jiaqi, et al.
Published: (2025)
Test-Time Low Rank Adaptation via Confidence Maximization for Zero-Shot Generalization of Vision-Language Models
by: Imam, Raza, et al.
Published: (2024)
by: Imam, Raza, et al.
Published: (2024)
AgMMU: A Comprehensive Agricultural Multimodal Understanding Benchmark
by: Gauba, Aruna, et al.
Published: (2025)
by: Gauba, Aruna, et al.
Published: (2025)
Vision-Language Navigation with Energy-Based Policy
by: Liu, Rui, et al.
Published: (2024)
by: Liu, Rui, et al.
Published: (2024)
Multi-Label Out-of-Distribution Detection with Spectral Normalized Joint Energy
by: Mei, Yihan, et al.
Published: (2024)
by: Mei, Yihan, et al.
Published: (2024)
NoPo-Avatar: Generalizable and Animatable Avatars from Sparse Inputs without Human Poses
by: Wen, Jing, et al.
Published: (2025)
by: Wen, Jing, et al.
Published: (2025)
LIFe-GoM: Generalizable Human Rendering with Learned Iterative Feedback Over Multi-Resolution Gaussians-on-Mesh
by: Wen, Jing, et al.
Published: (2025)
by: Wen, Jing, et al.
Published: (2025)
EnergyFormer: Energy Attention with Fourier Embedding for Hyperspectral Image Classification
by: Sohail, Saad, et al.
Published: (2025)
by: Sohail, Saad, et al.
Published: (2025)
Domain Adaptable Fine-Tune Distillation Framework For Advancing Farm Surveillance
by: Imam, Raza, et al.
Published: (2024)
by: Imam, Raza, et al.
Published: (2024)
Human-like Navigation in a World Built for Humans
by: Chandaka, Bhargav, et al.
Published: (2025)
by: Chandaka, Bhargav, et al.
Published: (2025)
Double-Exponential Increases in Inference Energy: The Cost of the Race for Accuracy
by: Yang, Zeyu, et al.
Published: (2024)
by: Yang, Zeyu, et al.
Published: (2024)
Energy-Latency Attacks via Sponge Poisoning
by: Cinà, Antonio Emanuele, et al.
Published: (2022)
by: Cinà, Antonio Emanuele, et al.
Published: (2022)
Towards Reproducible Learning-based Compression
by: Pang, Jiahao, et al.
Published: (2024)
by: Pang, Jiahao, et al.
Published: (2024)
Inducing High Energy-Latency of Large Vision-Language Models with Verbose Images
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
LidarDM: Generative LiDAR Simulation in a Generated World
by: Zyrianov, Vlas, et al.
Published: (2024)
by: Zyrianov, Vlas, et al.
Published: (2024)
Better, Stronger, Faster: Tackling the Trilemma in MLLM-based Segmentation with Simultaneous Textual Mask Prediction
by: Liu, Jiazhen, et al.
Published: (2025)
by: Liu, Jiazhen, et al.
Published: (2025)
TriAlignXA: An Explainable Trilemma Alignment Framework for Trustworthy Agri-product Grading
by: Xie, Jianfei, et al.
Published: (2025)
by: Xie, Jianfei, et al.
Published: (2025)
RepLDM: Reprogramming Pretrained Latent Diffusion Models for High-Quality, High-Efficiency, High-Resolution Image Generation
by: Cao, Boyuan, et al.
Published: (2024)
by: Cao, Boyuan, et al.
Published: (2024)
Energy-Latency Manipulation of Multi-modal Large Language Models via Verbose Samples
by: Gao, Kuofeng, et al.
Published: (2024)
by: Gao, Kuofeng, et al.
Published: (2024)
LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models
by: Yu, Haojie, et al.
Published: (2025)
by: Yu, Haojie, et al.
Published: (2025)
Micro-Structures Graph-Based Point Cloud Registration for Balancing Efficiency and Accuracy
by: Zhang, Rongling, et al.
Published: (2024)
by: Zhang, Rongling, et al.
Published: (2024)
Structure from Duplicates: Neural Inverse Graphics from a Pile of Objects
by: Cheng, Tianhang, et al.
Published: (2024)
by: Cheng, Tianhang, et al.
Published: (2024)
GoMAvatar: Efficient Animatable Human Modeling from Monocular Video Using Gaussians-on-Mesh
by: Wen, Jing, et al.
Published: (2024)
by: Wen, Jing, et al.
Published: (2024)
MonoPatchNeRF: Improving Neural Radiance Fields with Patch-based Monocular Guidance
by: Wu, Yuqun, et al.
Published: (2024)
by: Wu, Yuqun, et al.
Published: (2024)
Plenoptic PNG: Real-Time Neural Radiance Fields in 150 KB
by: Lee, Jae Yong, et al.
Published: (2024)
by: Lee, Jae Yong, et al.
Published: (2024)
Navigating the Accuracy-Size Trade-Off with Flexible Model Merging
by: Dhasade, Akash, et al.
Published: (2025)
by: Dhasade, Akash, et al.
Published: (2025)
WebForge: Breaking the Realism-Reproducibility-Scalability Trilemma in Browser Agent Benchmark
by: Yuan, Peng, et al.
Published: (2026)
by: Yuan, Peng, et al.
Published: (2026)
ObjectCompose: Evaluating Resilience of Vision-Based Models on Object-to-Background Compositional Changes
by: Malik, Hashmat Shadab, et al.
Published: (2024)
by: Malik, Hashmat Shadab, et al.
Published: (2024)
Enhancing Boundary Segmentation for Topological Accuracy with Skeleton-based Methods
by: Liu, Chuni, et al.
Published: (2024)
by: Liu, Chuni, et al.
Published: (2024)
DISHA: Low-Energy Sparse Transformer at Edge for Outdoor Navigation for the Visually Impaired Individuals
by: Nagil, Praveen, et al.
Published: (2024)
by: Nagil, Praveen, et al.
Published: (2024)
Towards Low-Latency Event Stream-based Visual Object Tracking: A Slow-Fast Approach
by: Wang, Shiao, et al.
Published: (2025)
by: Wang, Shiao, et al.
Published: (2025)
Improving Facial Landmark Detection Accuracy and Efficiency with Knowledge Distillation
by: Hong, Zong-Wei, et al.
Published: (2024)
by: Hong, Zong-Wei, et al.
Published: (2024)
EdgePoint2: Compact Descriptors for Superior Efficiency and Accuracy
by: Yao, Haodi, et al.
Published: (2025)
by: Yao, Haodi, et al.
Published: (2025)
WorldDreamer: Towards General World Models for Video Generation via Predicting Masked Tokens
by: Wang, Xiaofeng, et al.
Published: (2024)
by: Wang, Xiaofeng, et al.
Published: (2024)
AutoVFX: Physically Realistic Video Editing from Natural Language Instructions
by: Hsu, Hao-Yu, et al.
Published: (2024)
by: Hsu, Hao-Yu, et al.
Published: (2024)
Generalizable Sparse-View 3D Reconstruction from Unconstrained Images
by: Gupta, Vinayak, et al.
Published: (2026)
by: Gupta, Vinayak, et al.
Published: (2026)
AD-GS: Object-Aware B-Spline Gaussian Splatting for Self-Supervised Autonomous Driving
by: Xu, Jiawei, et al.
Published: (2025)
by: Xu, Jiawei, et al.
Published: (2025)
Similar Items
-
EFSA: Episodic Few-Shot Adaptation for Text-to-Image Retrieval
by: Huzaifa, Muhammad, et al.
Published: (2024) -
CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets
by: Amangeldi, Aidar, et al.
Published: (2025) -
PhysGen3D: Crafting a Miniature Interactive World from a Single Image
by: Chen, Boyuan, et al.
Published: (2025) -
Towards Latency-Aware 3D Streaming Perception for Autonomous Driving
by: Peng, Jiaqi, et al.
Published: (2025) -
Test-Time Low Rank Adaptation via Confidence Maximization for Zero-Shot Generalization of Vision-Language Models
by: Imam, Raza, et al.
Published: (2024)