ACM Transactions on Architecture and Code Optimization

Papers
(The median citation count of ACM Transactions on Architecture and Code Optimization is 1. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2021-11-01 to 2025-11-01.)
ArticleCitations
Performance, Energy and NVM Lifetime-Aware Data Structure Refinement and Placement for Heterogeneous Memory Systems44
An Intelligent Scheduling Approach on Mobile OS for Optimizing UI Smoothness and Power36
TNT: A Modular Approach to Traversing Physically Heterogeneous NOCs at Bare-wire Latency31
Object Intersection Captures on Interactive Apps to Drive a Crowd-sourced Replay-based Compiler Optimization26
Highly Efficient Self-checking Matrix Multiplication on Tiled AMX Accelerators26
ASM: An Adaptive Secure Multicore for Co-located Mutually Distrusting Processes22
TransCL: An Automatic CUDA-to-OpenCL Programs Transformation Framework21
Accelerating Verifiable Queries over Blockchain Database System Using Processing-in-memory20
ModNEF : An Open Source Modular Neuromorphic Emulator for FPGA for Low-Power In-Edge Artificial Intelligence20
Intra-request Lag-aware Cache Management to Enhance I/O Responsiveness of SSDs17
ESMPC: An Efficient Neural Network Training Framework for Secure Two- and Three-Party Computation17
Tiaozhuan: A General and Efficient Indirect Branch Optimization for Binary Translation16
An Accelerator for Sparse Convolutional Neural Networks Leveraging Systolic General Matrix-matrix Multiplication16
COER: A Network Interface Offloading Architecture for RDMA and Congestion Control Protocol Codesign15
DCMA: Accelerating Parallel DMA Transfers with a Multi-Port Direct Cached Memory Access in a Massive-Parallel Vector Processor14
Fast Convolution Meets Low Precision: Exploring Efficient Quantized Winograd Convolution on Modern CPUs13
Source Matching and Rewriting for MLIR Using String-Based Automata13
SIMD-Matcher: A SIMD-based Arbitrary Matching Framework13
A Concise Concurrent B + -Tree for Persistent Memory13
Locality-Aware CTA Scheduling for Gaming Applications12
FlashGEMM: Optimizing Sequences of Matrix Multiplication by Exploiting Data Reuse on CPUs12
Building a Fast and Efficient LSM-tree Store by Integrating Local Storage with Cloud Storage12
Mentor: A Memory-Efficient Sparse-dense Matrix Multiplication Accelerator Based on Column-Wise Product11
iSwap: A New Memory Page Swap Mechanism for Reducing Ineffective I/O Operations in Cloud Environments11
A NUMA-Aware Version of an Adaptive Self-Scheduling Loop Scheduler10
Accelerating Video Captioning on Heterogeneous System Architectures10
DeepZoning: Re-accelerate CNN Inference with Zoning Graph for Heterogeneous Edge Cluster9
COX : Exposing CUDA Warp-level Functions to CPUs9
MLKAPS: Machine Learning and Adaptive Sampling for HPC Kernel Auto-tuning9
AG-SpTRSV: An Automatic Framework to Optimize Sparse Triangular Solve on GPUs9
GraphSER: Distance-Aware Stream-Based Edge Repartition for Many-Core Systems9
Flexible and Effective Object Tiering for Heterogeneous Memory Systems9
SnsBooster: Enhancing Sampling-based μ Arch Evaluation Efficiency through Online Performance Sensitivity Analysis8
An FPGA Overlay for CNN Inference with Fine-grained Flexible Parallelism8
Accelerating Nearest Neighbor Search in 3D Point Cloud Registration on GPUs8
NEM-GNN: DAC/ADC-less, Scalable, Reconfigurable, Graph and Sparsity-Aware Near-Memory Accelerator for Graph Neural Networks8
Quantifying Resource Contention of Co-located Workloads with the System-level Entropy8
ODGS: Dependency-Aware Scheduling for High-Level Synthesis with Graph Neural Network and Reinforcement Learning8
Accelerating Parallel Structures in DNNs via Parallel Fusion and Operator Co-Optimization8
Efficient Cross-platform Multiplexing of Hardware Performance Counters via Adaptive Grouping7
Joint Program and Layout Transformations to Enable Convolutional Operators on Specialized Hardware Based on Constraint Programming7
Sectored DRAM: A Practical Energy-Efficient and High-Performance Fine-Grained DRAM Architecture7
A Fast and Flexible FPGA-based Accelerator for Natural Language Processing Neural Networks7
EXPERTISE: An Effective Software-level Redundant Multithreading Scheme against Hardware Faults7
Advancing Direct Convolution Using Convolution Slicing Optimization and ISA Extensions7
BridgeGC: An Efficient Cross-Level Garbage Collector for Big Data Frameworks7
PowerMorph: QoS-Aware Server Power Reshaping for Data Center Regulation Service6
Environmental Condition Aware Super-Resolution Acceleration Framework in Server-Client Hierarchies6
RT-GNN: Accelerating Sparse Graph Neural Networks by Tensor-CUDA Kernel Fusion6
Orchard: Heterogeneous Parallelism and Fine-grained Fusion for Complex Tree Traversals6
RaNAS: Resource-Aware Neural Architecture Search for Edge Computing6
Towards High Performance QNNs via Distribution-Based CNOT Gate Reduction6
HEngine: A High Performance Optimization Framework on a GPU for Homomorphic Encryption6
MemoriaNova: Optimizing Memory-Aware Model Inference for Edge Computing6
HyGain: High-performance, Energy-efficient Hybrid Gain Cell-based Cache Hierarchy6
TPRepair: Tree-based Pipelined Repair in Clustered Storage Systems6
Low-power Near-data Instruction Execution Leveraging Opcode-based Timing Analysis6
Multi-objective Hardware-aware Neural Architecture Search with Pareto Rank-preserving Surrogate Models6
DTAP: Accelerating Strongly-Typed Programs with Data Type-Aware Hardware Prefetching6
Accelerating Convolutional Neural Network by Exploiting Sparsity on GPUs5
Stripe-schedule Aware Repair in Erasure-coded Clusters with Heterogeneous Star Networks5
WIPE: A Write-Optimized Learned Index for Persistent Memory5
ERASE: Energy Efficient Task Mapping and Resource Management for Work Stealing Runtimes5
gECC: A GPU-based high-throughput framework for Elliptic Curve Cryptography5
Mobile-3DCNN: An Acceleration Framework for Ultra-Real-Time Execution of Large 3D CNNs on Mobile Devices5
GraphTune: An Efficient Dependency-Aware Substrate to Alleviate Irregularity in Concurrent Graph Processing5
RACER: Avoiding End-to-End Slowdowns in Accelerated Chip Multi-Processors5
Toward Comprehensive Design Space Exploration on Heterogeneous Multi-core Processors5
SimTrace: Exploiting Spatial and Temporal Sampling for Large-Scale Performance Analysis5
EDAS: Enabling Fast Data Loading for GPU Serverless Computing5
FlexHM: A Practical System for Heterogeneous Memory with Flexible and Efficient Performance Optimizations5
CGCGraph: Efficient CPU-GPU Co-execution for Concurrent Dynamic Graph Processing5
Improving Utilization of Dataflow Unit for Multi-Batch Processing5
OptiFX: Automatic Optimization for Convolutional Neural Networks with Aggressive Operator Fusion on GPUs5
A Stable Idle Time Detection Platform for Real I/O Workloads5
Towards Optimizing Learned Index for High Performance, Memory Efficiency and NUMA Awareness5
Exploring Data Layout for Sparse Tensor Times Dense Matrix on GPUs5
Koala: Efficient Pipeline Training through Automated Schedule Searching on Domain-Specific Language4
CoolDC: A Cost-Effective Immersion-Cooled Datacenter with Workload-Aware Temperature Scaling4
Efficient Flexible Edge Inference for Mixed-Precision Quantized DNN using Customized RISC-V Core4
BullsEye : Scalable and Accurate Approximation Framework for Cache Miss Calculation4
Capability-Based Efficient Data Transmission Mechanism for Serverless Computing4
Shift-CIM: In-SRAM Alignment To Support General-Purpose Bit-level Sparsity Exploration in SRAM Multiplication4
Address/Data Instruction Steering in Clustered General Purpose Processors4
JiuJITsu: Removing Gadgets with Safe Register Allocation for JIT Code Generation4
Architectural Support for Sharing, Isolating and Virtualizing FPGA Resources4
MetaEC: An Efficient and Resilient Erasure-Coded KV Store on Disaggregated Memory4
TSN Cache: Exploiting Data Localities in Graph Computing Applications4
x Meta : SSD-HDD-hybrid Optimization for Metadata Maintenance of Cloud-scale Object Storage4
Scale-out Systolic Arrays4
CASHT: Contention Analysis in Shared Hierarchies with Thefts4
Architecting Optically Controlled Phase Change Memory4
SplitZNS: Towards an Efficient LSM-Tree on Zoned Namespace SSDs4
SAL: Optimizing the Dataflow of Spin-based Architectures for Lightweight Neural Networks3
Compressing and Accelerating Sparse CNNs Using Sign-Reserved Toeplitz Filters and Input Activation Density-aware Dataflow3
Towards Efficient Extendible Perfect Hashing for Hybrid PM-DRAM Memory3
Asynchronous Memory Access Unit: Exploiting Massive Parallelism for Far Memory Access3
Preserving Addressability Upon GC-Triggered Data Movements on Non-Volatile Memory3
Iterating Pointers: Enabling Static Analysis for Loop-based Pointers3
MicroProf : Code-level Attribution of Unnecessary Data Transfer in Microservice Applications3
CoNST: Code Generator for Sparse Tensor Networks3
Jointly Optimizing Job Assignment and Resource Partitioning for Improving System Throughput in Cloud Datacenters3
E-BATCH: Energy-Efficient and High-Throughput RNN Batching3
A Low-latency On-chip Cache Hierarchy for Load-to-use Stall Reduction in GPUs3
Cheetah: Accelerating Dynamic Graph Mining with Grouping Updates3
FlowPix: Accelerating Image Processing Pipelines on an FPGA Overlay using a Domain Specific Compiler3
Optimizing OpenCL Barrier Synchronization and Memory Efficiency on Multi-Core DSPs3
An Example of Parallel Merkle Tree Traversal: Post-Quantum Leighton-Micali Signature on the GPU3
RaKV: A Write-Optimized LSM Store for Cloud Block Storage with Robust SLA3
TLB-pilot: Mitigating TLB Contention Attack on GPUs with Microarchitecture-Aware Scheduling3
PARALiA: A Performance Aware Runtime for Auto-tuning Linear Algebra on Heterogeneous Systems3
PANDA: Adaptive Prefetching and Decentralized Scheduling for Dataflow Architectures3
High-performance Deterministic Concurrency Using Lingua Franca3
Memory-Aware Functional IR for Higher-Level Synthesis of Accelerators3
Matrix: Multi-Cipher Structures Dataflow for Parallel and Pipelined TFHE Accelerator3
Abakus: Accelerating k -mer Counting with Storage Technology3
Consequence-based Clustered Architecture3
3D GNLM: Efficient 3D Non-Local Means Kernel with Nested Reuse Strategies for Embedded GPUs3
An FPGA-based Approach to Evaluate Thermal and Resource Management Strategies of Many-core Processors3
In-SRAM Parallel Data Shuffle2
Cerberus: Triple Mode Acceleration of Sparse Matrix and Vector Multiplication2
Bubble-Swap Flow Control2
SuccinctKV: a CPU-efficient LSM-tree Based KV Store with Scan-based Compaction2
ShieldCXL: A Practical Obliviousness Support with Sealed CXL Memory2
Delay-on-Squash: Stopping Microarchitectural Replay Attacks in Their Tracks2
QuCloud+: A Holistic Qubit Mapping Scheme for Single/Multi-programming on 2D/3D NISQ Quantum Computers2
MemHC: An Optimized GPU Memory Management Framework for Accelerating Many-body Correlation2
PARADISE: Criticality-Aware Instruction Reordering for Power Attack Resistance2
The Impact of Page Size and Microarchitecture on Instruction Address Translation Overhead2
Design and Implementation for Nonblocking Execution in GraphBLAS: Tradeoffs and Performance2
A Pressure-Aware Policy for Contention Minimization on Multicore Systems2
Conflict Management in Vector Register Files2
SMT-Based Contention-Free Task Mapping and Scheduling on 2D/3D SMART NoC with Mixed Dimension-Order Routing2
Winols: A Large-Tiling Sparse Winograd CNN Accelerator on FPGAs2
Understanding Silent Data Corruption in Processors for Mitigating its Effects2
SPIRIT: Scalable and Persistent In-Memory Indices for Real-Time Search2
ZNSFQ: An Efficient and High-Performance Fair Queue Scheduling Scheme for ZNS SSDs2
An Optimized GPU Implementation for GIST Descriptor2
At the Locus of Performance: Quantifying the Effects of Copious 3D-Stacked Cache on HPC Workloads2
ReSA: Reconfigurable Systolic Array for Multiple Tiny DNN Tensors2
SpecTerminator: Blocking Speculative Side Channels Based on Instruction Classes on RISC-V2
GenCNN: A Partition-Aware Multi-Objective Mapping Framework for CNN Accelerators Based on Genetic Algorithm2
A Case For Intra-rack Resource Disaggregation in HPC2
A Data-Loader Tunable Knob to Shorten GPU Idleness for Distributed Deep Learning2
Register-Pressure-Aware Instruction Scheduling Using Ant Colony Optimization2
Puppeteer: A Random Forest Based Manager for Hardware Prefetchers Across the Memory Hierarchy2
Supporting Dynamic Program Sizes in Deep Learning-Based Cost Models for Code Optimization2
Compiler Support for Sparse Tensor Computations in MLIR2
GPU Domain Specialization via Composable On-Package Architecture2
ReIPE: Recycling Idle PEs in CNN Accelerator for Vulnerable Filters Soft-Error Detection2
Approx-RM: Reducing Energy on Heterogeneous Multicore Processors under Accuracy and Timing Constraints2
SSD-SGD: Communication Sparsification for Distributed Deep Learning Training2
HAVIT: An Efficient Hardware-Accelerator for Vision Transformer with Informative Patch Selection Techniques2
PIMSAB: A Processing-In-Memory System with Spatially-Aware Communication and Bit-Serial-Aware Computation2
HAIR: Halving the Area of the Integer Register File with Odd/Even Banking2
Data Deduplication Based on Content Locality of Transactions to Enhance Blockchain Scalability2
The Forward Slice Core: A High-Performance, Yet Low-Complexity Microarchitecture2
ApSpGEMM: Accelerating Large-scale SpGEMM with Heterogeneous Collaboration and Adaptive Panel2
GraphService: Topology-aware Constructor for Large-scale Graph Applications2
Optimization of Sparse Matrix Computation for Algebraic Multigrid on GPUs2
SAC: An Ultra-Efficient Spin-based Architecture for Compressed DNNs2
HotLD: a Workload-Aware Method for Global Code-Layout Optimization of Shared Libraries1
TianheGraph: Topology-aware Graph Processing1
ApHMM: Accelerating Profile Hidden Markov Models for Fast and Energy-efficient Genome Analysis1
DFGAS: Exploring the Balance of HW-SW Scheduling through the DFG-Aware Scheme1
The Design of an Efficient Lossy Compressor for Time Series Databases1
Second-level Caches: Not for Instructions1
ShuffleInfer: Disaggregate LLM Inference for Mixed Downstream Workloads1
GOLDYLOC: Global Optimizations & Lightweight Dynamic Logic for Concurrency1
Unveiling and Evaluating Vulnerabilities in Branch Predictors via a Three-Step Modeling Methodology1
GiantVM: A Novel Distributed Hypervisor for Resource Aggregation with DSM-aware Optimizations1
HAKV: A Hotness-Aware Zone Management Approach to Optimizing Performance of LSM-tree-based Key-Value Stores1
PRAGA: A Priority-Aware Hardware/Software Co-design for High-Throughput Graph Processing Acceleration1
A 2 : Towards Accelerator Level Parallelism for Autonomous Micromobility Systems1
A Survey of General-purpose Polyhedral Compilers1
Optimizing Garbage Collection for ZNS SSDs via In-storage Data Migration and Address Remapping1
Scheduling Language Chronology: Past, Present, and Future1
PDGNN: Efficient Micro-batch GNN Training via Degree-Pruned Partitioning and Redundancy Elimination1
Solving Sparse Assignment Problems on FPGAs1
Gator: Accelerating Graph Attention Networks by Jointly Optimizing Attention and Graph Processing1
DCSolver: Accelerating Sparse Iterative Solvers via Divide-and-Conquer on GPUs1
Partitioned Scheduling and Analysis for a Typed DAG Task on Heterogeneous Multi-Cores1
A Lock-free RDMA-friendly Index in CPU-parsimonious Environments1
PiDRAM: A Holistic End-to-end FPGA-based Framework for Processing-in-DRAM1
Assessing the Impact of Compiler Optimizations on GPUs Reliability1
Critical Data Backup with Hybrid Flash-Based Consumer Devices1
DELTA: Memory-Efficient Training via Dynamic Fine-Grained Recomputation and Swapping1
FlexPointer: Fast Address Translation Based on Range TLB and Tagged Pointers1
TEA+ : A Novel Temporal Graph Random Walk Engine with Hybrid Storage Architecture1
VersaTile: Flexible Tiled Architectures via Associative Processors1
Symbolic Analysis for Data Plane Programs Specialization1
The Droplet Search Algorithm for Kernel Scheduling1
Overlapping Aware Data Placement Optimizations for LSM Tree-Based Store on ZNS SSDs1
Lock-Free High-performance Hashing for Persistent Memory via PM-aware Holistic Optimization1
IBing: An Efficient Interleaved Bidirectional Ring All-Reduce Algorithm for Gradient Synchronization1
Hardware-hardened Sandbox Enclaves for Trusted Serverless Computing1
MetaSys: A Practical Open-source Metadata Management System to Implement and Evaluate Cross-layer Optimizations1
Constructing a Supplementary Benchmark Suite to Represent Android Applications with User Interactions by using Performance Counters1
CARL: Compiler Assigned Reference Leasing1
Fast One-Sided RDMA-Based State Machine Replication for Disaggregated Memory1
A Sparsity-Aware Autonomous Path Planning Accelerator with HW/SW Co-Design and Multi-Level Dataflow Optimization1
Turn-based Spatiotemporal Coherence for GPUs1
Unleashing Parallelism with Elastic-Barriers1
ScaleGS: Closing the Gap between Real-time 3D Gaussian Splatting and Real-time XR Rendering1
gHyPart: GPU-friendly End-to-End Hypergraph Partitioner1
MUA-Router: Maximizing the Utility-of-Allocation for On-chip Pipelining Routers1
JUNO++: Optimizing ANNS and Enabling Efficient Sparse Attention in LLM via Ray Tracing Core1
Mapi-Pro: An Energy Efficient Memory Mapping Technique for Intermittent Computing1
LitTLS: Lightweight Thread-Level Speculation on Little Cores1
0.50055503845215