FALCON: Feedback-driven Adaptive Long/short-term memory reinforced Coding Optimization system
Fuente:
arXiv
Saved in:
| Main Authors: | Li, Zeyuan, He, Yangfan, He, Lewei, Wang, Jianhui, Shi, Tianyu, Lei, Bin, Li, Yuchen, Chen, Qiuwu |
|---|---|
| Format: | Preprint |
| Published: |
2024
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
AutoCoder: Enhancing Code Large Language Model with \textsc{AIEV-Instruct}
by: Lei, Bin, et al.
Published: (2024)
by: Lei, Bin, et al.
Published: (2024)
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
by: Zhou, Fang, et al.
Published: (2026)
by: Zhou, Fang, et al.
Published: (2026)
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
by: Wang, Yuxin, et al.
Published: (2023)
by: Wang, Yuxin, et al.
Published: (2023)
Attributing the System's Overall Effect to its Components
by: Wang, Chenxi, et al.
Published: (2026)
by: Wang, Chenxi, et al.
Published: (2026)
Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
by: Ren, Jie, et al.
Published: (2025)
by: Ren, Jie, et al.
Published: (2025)
LiveTune: Dynamic Parameter Tuning for Feedback-Driven Optimization
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2023)
by: Shabgahi, Soheil Zibakhsh, et al.
Published: (2023)
Inference performance evaluation for LLMs on edge devices with a novel benchmarking framework and metric
by: Chen, Hao, et al.
Published: (2025)
by: Chen, Hao, et al.
Published: (2025)
Mosaic: Cross-Modal Clustering for Efficient Video Understanding
by: Wang, Tuowei, et al.
Published: (2026)
by: Wang, Tuowei, et al.
Published: (2026)
From Profiling to Optimization: Unveiling the Profile Guided Optimization
by: Liu, Bingxin, et al.
Published: (2025)
by: Liu, Bingxin, et al.
Published: (2025)
Personalized Model-Based Design of Human Centric AI enabled CPS for Long term usage
by: Ngabonziza, Bernard, et al.
Published: (2026)
by: Ngabonziza, Bernard, et al.
Published: (2026)
CesASMe and Staticdeps: static detection of memory-carried dependencies for code analyzers
by: Bastian, Théophile, et al.
Published: (2024)
by: Bastian, Théophile, et al.
Published: (2024)
Achieving Consistent and Comparable CPU Evaluation
by: Wang, Chenxi, et al.
Published: (2024)
by: Wang, Chenxi, et al.
Published: (2024)
Optimizing Winograd Convolution on ARMv8 processors
by: Gui, Haoyuan, et al.
Published: (2024)
by: Gui, Haoyuan, et al.
Published: (2024)
Entanglement buffering with two quantum memories
by: Davies, Bethany, et al.
Published: (2023)
by: Davies, Bethany, et al.
Published: (2023)
Entanglement buffering with multiple quantum memories
by: Iñesta, Álvaro G., et al.
Published: (2025)
by: Iñesta, Álvaro G., et al.
Published: (2025)
AGIPC: Adaptive In-Solve Algebraic Coarsening for GPU IPC
by: Wang, Xuan, et al.
Published: (2026)
by: Wang, Xuan, et al.
Published: (2026)
DSO: A GPU Energy Efficiency Optimizer by Fusing Dynamic and Static Information
by: Wang, Qiang, et al.
Published: (2024)
by: Wang, Qiang, et al.
Published: (2024)
H2EAL: Hybrid-Bonding Architecture with Hybrid Sparse Attention for Efficient Long-Context LLM Inference
by: Fu, Zizhuo, et al.
Published: (2025)
by: Fu, Zizhuo, et al.
Published: (2025)
DRIM-ANN: An Approximate Nearest Neighbor Search Engine based on Commercial DRAM-PIMs
by: Chen, Mingkai, et al.
Published: (2024)
by: Chen, Mingkai, et al.
Published: (2024)
Spatiotemporal Non-Uniformity-Aware Online Task Scheduling in Collaborative Edge Computing for Industrial Internet of Things
by: Li, Yang, et al.
Published: (2025)
by: Li, Yang, et al.
Published: (2025)
Artifacts and bodily acts of memory: art and life to think the fine line of the present
by: Barretto, Eleonora Frenkel
Published: (2023)
by: Barretto, Eleonora Frenkel
Published: (2023)
How Much Parallelism Is "Free"? A Principle of Near-Free Parallelism for Parallel Decoding
by: He, Minghua, et al.
Published: (2026)
by: He, Minghua, et al.
Published: (2026)
Scalable Packed Layouts for Vector-Length-Agnostic ML Code Generation
by: Beysel, Ege, et al.
Published: (2026)
by: Beysel, Ege, et al.
Published: (2026)
Can We Make Code Green? Understanding Trade-Offs in LLMs vs. Human Code Optimizations
by: Rani, Pooja, et al.
Published: (2025)
by: Rani, Pooja, et al.
Published: (2025)
Can Increasing the Hit Ratio Hurt Cache Throughput? (Long Version)
by: Qiu, Ziyue, et al.
Published: (2024)
by: Qiu, Ziyue, et al.
Published: (2024)
Tracing Optimization for Performance Modeling and Regression Detection
by: Shahedi, Kaveh, et al.
Published: (2024)
by: Shahedi, Kaveh, et al.
Published: (2024)
Flex Attention: A Programming Model for Generating Optimized Attention Kernels
by: Dong, Juechu, et al.
Published: (2024)
by: Dong, Juechu, et al.
Published: (2024)
Systematic Evaluation of Optimization Techniques for Long-Context Language Models
by: Ahmed, Ammar, et al.
Published: (2025)
by: Ahmed, Ammar, et al.
Published: (2025)
Ecoscape: Fault Tolerance Benchmark for Adaptive Remediation Strategies in Real-Time Edge ML
by: Reiter, Hendrik, et al.
Published: (2025)
by: Reiter, Hendrik, et al.
Published: (2025)
ARKV: Adaptive and Resource-Efficient KV Cache Management under Limited Memory Budget for Long-Context Inference in LLMs
by: Lei, Jianlong, et al.
Published: (2026)
by: Lei, Jianlong, et al.
Published: (2026)
Recurrent CircuitSAT Sampling for Sequential Circuits
by: Ardakani, Arash, et al.
Published: (2025)
by: Ardakani, Arash, et al.
Published: (2025)
Two Criteria for Performance Analysis of Optimization Algorithms
by: Jing, Yunpeng, et al.
Published: (2024)
by: Jing, Yunpeng, et al.
Published: (2024)
Performance Characterization and Optimizations of Traditional ML Applications
by: Kumar, Harsh, et al.
Published: (2024)
by: Kumar, Harsh, et al.
Published: (2024)
A Zoned Storage Optimized Flash Cache on ZNS SSDs
by: Yang, Chongzhuo, et al.
Published: (2024)
by: Yang, Chongzhuo, et al.
Published: (2024)
On the Design of Capacity-Achieving Distributions for Discrete-Time Poisson Channel with Low-Precision ADCs
by: Li, Qianqian, et al.
Published: (2025)
by: Li, Qianqian, et al.
Published: (2025)
Caspar: CUDA Accelerator for Symbolic Programming with Adaptive Reordering
by: Martens, Emil, et al.
Published: (2026)
by: Martens, Emil, et al.
Published: (2026)
Should AI Optimize Your Code? A Comparative Study of Classical Optimizing Compilers Versus Current Large Language Models
by: Rosas, Miguel Romero, et al.
Published: (2024)
by: Rosas, Miguel Romero, et al.
Published: (2024)
Opal: A Modular Framework for Optimizing Performance using Analytics and LLMs
by: Zaeed, Mohammad, et al.
Published: (2025)
by: Zaeed, Mohammad, et al.
Published: (2025)
Performance Optimization of 3D Stencil Computation on ARM Scalable Vector Extension
by: Chen, Hongguang
Published: (2025)
by: Chen, Hongguang
Published: (2025)
Dual-Signal Adaptive KV-Cache Optimization for Long-Form Video Understanding in Vision-Language Models
by: Sai, Vishnu, et al.
Published: (2026)
by: Sai, Vishnu, et al.
Published: (2026)
Similar Items
-
AutoCoder: Enhancing Code Large Language Model with \textsc{AIEV-Instruct}
by: Lei, Bin, et al.
Published: (2024) -
Long-term Monitoring of Kernel and Hardware Events to Understand Latency Variance
by: Zhou, Fang, et al.
Published: (2026) -
Fault-Tolerant Hybrid-Parallel Training at Scale with Reliable and Efficient In-memory Checkpointing
by: Wang, Yuxin, et al.
Published: (2023) -
Attributing the System's Overall Effect to its Components
by: Wang, Chenxi, et al.
Published: (2026) -
Machine Learning-Guided Memory Optimization for DLRM Inference on Tiered Memory
by: Ren, Jie, et al.
Published: (2025)