AVERY: Intent-Driven Adaptive VLM Split Computing via Embodied Self-Awareness for Efficient Disaster Response Systems
Fuente:
arXiv
Salvato in:
| Autori principali: | Bhattacharjya, Rajat, Wu, Sing-Yao, Oh, Hyunwoo, Nam, Chaewon, Koo, Suyeon, Imani, Mohsen, Bozorgzadeh, Elaheh, Dutt, Nikil |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
HYPERDOA: Robust and Efficient DoA Estimation using Hyperdimensional Computing
di: Bhattacharjya, Rajat, et al.
Pubblicazione: (2025)
di: Bhattacharjya, Rajat, et al.
Pubblicazione: (2025)
T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU Reorganization
di: Oh, Hyunwoo, et al.
Pubblicazione: (2025)
di: Oh, Hyunwoo, et al.
Pubblicazione: (2025)
MUSIC-lite: Efficient MUSIC using Approximate Computing: An OFDM Radar Case Study
di: Bhattacharjya, Rajat, et al.
Pubblicazione: (2024)
di: Bhattacharjya, Rajat, et al.
Pubblicazione: (2024)
ACCESS-AV: Adaptive Communication-Computation Codesign for Sustainable Autonomous Vehicle Localization in Smart Factories
di: Bhattacharjya, Rajat, et al.
Pubblicazione: (2025)
di: Bhattacharjya, Rajat, et al.
Pubblicazione: (2025)
MOFCO: Mobility- and Migration-Aware Task Offloading in Three-Layer Fog Computing Environments
di: Mahdizadeh, Soheil, et al.
Pubblicazione: (2025)
di: Mahdizadeh, Soheil, et al.
Pubblicazione: (2025)
CRAFT: Latency and Cost-Aware Genetic-Based Framework for Node Placement in Edge-Fog Environments
di: Mahdizadeh, Soheil, et al.
Pubblicazione: (2025)
di: Mahdizadeh, Soheil, et al.
Pubblicazione: (2025)
Sequence-Aware Split Heuristic to Mitigate SM Underutilization in FlashAttention-3 Low-Head-Count Decoding
di: Font, Martí Llopart, et al.
Pubblicazione: (2026)
di: Font, Martí Llopart, et al.
Pubblicazione: (2026)
QUILL: An Algorithm-Architecture Co-Design for Cache-Local Deformable Attention
di: Oh, Hyunwoo, et al.
Pubblicazione: (2025)
di: Oh, Hyunwoo, et al.
Pubblicazione: (2025)
TorR: Towards Brain-Inspired Task-Oriented Reasoning via Cache-Oriented Algorithm-Architecture Co-design
di: Oh, Hyunwoo, et al.
Pubblicazione: (2026)
di: Oh, Hyunwoo, et al.
Pubblicazione: (2026)
FPGA Innovation Research in the Netherlands: Present Landscape and Future Outlook
di: Alachiotis, Nikolaos, et al.
Pubblicazione: (2025)
di: Alachiotis, Nikolaos, et al.
Pubblicazione: (2025)
Intent-Driven Storage Systems: From Low-Level Tuning to High-Level Understanding
di: Bergman, Shai, et al.
Pubblicazione: (2025)
di: Bergman, Shai, et al.
Pubblicazione: (2025)
Analysing Mechanisms for Virtual Channel Management in Low-Diameter networks
di: Cano, Alejandro, et al.
Pubblicazione: (2023)
di: Cano, Alejandro, et al.
Pubblicazione: (2023)
RailX: A Flexible, Scalable, and Low-Cost Network Architecture for Hyper-Scale LLM Training Systems
di: Feng, Yinxiao, et al.
Pubblicazione: (2025)
di: Feng, Yinxiao, et al.
Pubblicazione: (2025)
SCENIC: Stream Computation-Enhanced SmartNIC
di: Ramhorst, Benjamin, et al.
Pubblicazione: (2026)
di: Ramhorst, Benjamin, et al.
Pubblicazione: (2026)
LACIN: Linearly Arranged Complete Interconnection Networks
di: Beivide, Ramón, et al.
Pubblicazione: (2026)
di: Beivide, Ramón, et al.
Pubblicazione: (2026)
cMPI: Using CXL Memory Sharing for MPI One-Sided and Two-Sided Inter-Node Communications
di: Wang, Xi, et al.
Pubblicazione: (2025)
di: Wang, Xi, et al.
Pubblicazione: (2025)
TT-Edge: A Hardware-Software Co-Design for Energy-Efficient Tensor-Train Decomposition on Edge AI
di: Kwak, Hyunseok, et al.
Pubblicazione: (2025)
di: Kwak, Hyunseok, et al.
Pubblicazione: (2025)
The DEEP-ER project: I/O and resiliency extensions for the Cluster-Booster architecture
di: Kreuzer, Anke, et al.
Pubblicazione: (2019)
di: Kreuzer, Anke, et al.
Pubblicazione: (2019)
An Evaluation and Comparison of GPU Hardware and Solver Libraries for Accelerating the OPM Flow Reservoir Simulator
di: Qiu, Tong Dong, et al.
Pubblicazione: (2023)
di: Qiu, Tong Dong, et al.
Pubblicazione: (2023)
SLIM: A Heterogeneous Accelerator for Edge Inference of Sparse Large Language Model via Adaptive Thresholding
di: Xu, Weihong, et al.
Pubblicazione: (2025)
di: Xu, Weihong, et al.
Pubblicazione: (2025)
COMET: A Framework for Modeling Compound Operation Dataflows with Explicit Collectives
di: Negi, Shubham, et al.
Pubblicazione: (2025)
di: Negi, Shubham, et al.
Pubblicazione: (2025)
MVDRAM: Enabling GeMV Execution in Unmodified DRAM for Low-Bit LLM Acceleration
di: Kubo, Tatsuya, et al.
Pubblicazione: (2025)
di: Kubo, Tatsuya, et al.
Pubblicazione: (2025)
PAM: Processing Across Memory Hierarchy for Efficient KV-centric LLM Serving System
di: Liu, Lian, et al.
Pubblicazione: (2026)
di: Liu, Lian, et al.
Pubblicazione: (2026)
RAPID-Graph: Recursive All-Pairs Shortest Paths Using Processing-in-Memory for Dynamic Programming on Graphs
di: Chen, Yanru, et al.
Pubblicazione: (2025)
di: Chen, Yanru, et al.
Pubblicazione: (2025)
Chopper: A Multi-Level GPU Characterization Tool & Derived Insights Into LLM Training Inefficiency
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
di: Kurzynski, Marco, et al.
Pubblicazione: (2025)
Towards Compute-Aware In-Switch Computing for LLMs Tensor-Parallelism on Multi-GPU Systems
di: Zhang, Chen, et al.
Pubblicazione: (2026)
di: Zhang, Chen, et al.
Pubblicazione: (2026)
Efficient deadlock avoidance for 2D mesh NoCs that use OQ or VOQ routers
di: Papaphilippou, Philippos, et al.
Pubblicazione: (2023)
di: Papaphilippou, Philippos, et al.
Pubblicazione: (2023)
DCRA: A Distributed Chiplet-based Reconfigurable Architecture for Irregular Applications
di: Orenes-Vera, Marcelo, et al.
Pubblicazione: (2023)
di: Orenes-Vera, Marcelo, et al.
Pubblicazione: (2023)
LFOC: A Lightweight Fairness-Oriented Cache Clustering Policy for Commodity Multicores
di: García-García, Adrián, et al.
Pubblicazione: (2024)
di: García-García, Adrián, et al.
Pubblicazione: (2024)
FlexVector: A SpMM Vector Processor with Flexible VRF for GCNs on Varying-Sparsity Graphs
di: Li, Bohan, et al.
Pubblicazione: (2026)
di: Li, Bohan, et al.
Pubblicazione: (2026)
iHAC: A Hybrid Cluster Architecture for Enhanced Performance and Resilience
di: Muntaka, Siddique Abubakr, et al.
Pubblicazione: (2026)
di: Muntaka, Siddique Abubakr, et al.
Pubblicazione: (2026)
NetSmith: An Optimization Framework for Machine-Discovered Network Topologies
di: Green, Conor, et al.
Pubblicazione: (2024)
di: Green, Conor, et al.
Pubblicazione: (2024)
SpArch: Efficient Architecture for Sparse Matrix Multiplication
di: Zhang, Zhekai, et al.
Pubblicazione: (2020)
di: Zhang, Zhekai, et al.
Pubblicazione: (2020)
Navigating the Landscape of Distributed File Systems: Architectures, Implementations, and Considerations
di: Pan, Xueting, et al.
Pubblicazione: (2024)
di: Pan, Xueting, et al.
Pubblicazione: (2024)
Knowledge-Guided Attention-Inspired Learning for Task Offloading in Vehicle Edge Computing
di: Ma, Ke, et al.
Pubblicazione: (2025)
di: Ma, Ke, et al.
Pubblicazione: (2025)
Efficient MoE Serving in the Memory-Bound Regime: Balance Activated Experts, Not Tokens
di: Yu, Yanpeng, et al.
Pubblicazione: (2025)
di: Yu, Yanpeng, et al.
Pubblicazione: (2025)
Design in Tiles: Automating GEMM Deployment on Tile-Based Many-PE Accelerators
di: Shen, Aofeng, et al.
Pubblicazione: (2025)
di: Shen, Aofeng, et al.
Pubblicazione: (2025)
FengHuang: Next-Generation Memory Orchestration for AI Inferencing
di: Li, Jiamin, et al.
Pubblicazione: (2025)
di: Li, Jiamin, et al.
Pubblicazione: (2025)
Evaluating Rapid Makespan Predictions for Heterogeneous Systems with Programmable Logic
di: Wilhelm, Martin, et al.
Pubblicazione: (2025)
di: Wilhelm, Martin, et al.
Pubblicazione: (2025)
Context-aware Simopt-Power: Using structural data with simulation metadata to optimise FPGA designs
di: Wadhwa, Eashan, et al.
Pubblicazione: (2026)
di: Wadhwa, Eashan, et al.
Pubblicazione: (2026)
Documenti analoghi
-
HYPERDOA: Robust and Efficient DoA Estimation using Hyperdimensional Computing
di: Bhattacharjya, Rajat, et al.
Pubblicazione: (2025) -
T-SAR: A Full-Stack Co-design for CPU-Only Ternary LLM Inference via In-Place SIMD ALU Reorganization
di: Oh, Hyunwoo, et al.
Pubblicazione: (2025) -
MUSIC-lite: Efficient MUSIC using Approximate Computing: An OFDM Radar Case Study
di: Bhattacharjya, Rajat, et al.
Pubblicazione: (2024) -
ACCESS-AV: Adaptive Communication-Computation Codesign for Sustainable Autonomous Vehicle Localization in Smart Factories
di: Bhattacharjya, Rajat, et al.
Pubblicazione: (2025) -
MOFCO: Mobility- and Migration-Aware Task Offloading in Three-Layer Fog Computing Environments
di: Mahdizadeh, Soheil, et al.
Pubblicazione: (2025)