IEEE Computer Architecture Letters

Papers
(The median citation count of IEEE Computer Architecture Letters is 1. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Toward Practical 128-Bit General Purpose Microarchitectures112
Characterization and Analysis of Text-to-Image Diffusion Models37
Exploration of Algorithm-Hardware Co-Design for Floating-Point Digital Compute-in-Memory27
Old is Gold: Optimizing Single-Threaded Applications With ExGen-Malloc25
Accelerating Programmable Bootstrapping Targeting Contemporary GPU Microarchitecture23
A Characterization of Generative Recommendation Models: Study of Hierarchical Sequential Transduction Unit22
NeuroMTA: Programmable Simulation Framework for Multi-Tile NPU Architectures19
The Architectural Sustainability Indicator15
SCALES: SCALable and Area-Efficient Systolic Accelerator for Ternary Polynomial Multiplication13
Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure12
De-Quantization Penalties for Interactive LLM Inference on Prosumer GPUs12
Context-Aware Set Dueling for Dynamic Policy Arbitration12
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference11
A Quantitative Analysis of Mamba-2-Based Large Language Model: Study of State Space Duality11
In-Depth Characterization of Machine Learning on an Optimized Multi-Party Computing Library10
OASIS: Outlier-Aware KV Cache Clustering for Scaling LLM Inference in CXL Memory Systems9
Time Series Machine Learning Models for Precise SSD Access Latency Prediction9
Improving Energy-Efficiency of Capsule Networks on Modern GPUs9
SoCurity: A Design Approach for Enhancing SoC Security9
AiDE: Attention-FFN Disaggregated Execution for Cost-Effective LLM Decoding on CXL-PNM8
Exploring KV Cache Quantization in Multimodal Large Language Model Inference8
Wafer-Scale GPU Memory Pool With In-Package Optics for Enhanced Capacity and Bandwidth8
RouteReplies: Alleviating Long Latency in Many-Chip-Module GPUs7
Straw: A Stress-Aware WL-Based Read Reclaim Technique for High-Density NAND Flash-Based SSDs7
A Case for In-Memory Random Scatter-Gather for Fast Graph Processing7
A Flexible Embedding-Aware Near Memory Processing Architecture for Recommendation System7
In-Memory Versioning (IMV)7
REDIT: Redirection-Enabled Memory-Side Directory Architecture for CXL Memory Fabric7
Disaggregated Speculative Decoding for Carbon-Efficient LLM Serving6
Reducing Metadata and Page Migration Overheads in CXL-Based Secure Tiered Memory6
Security Helper Chiplets: A New Paradigm for Secure Hardware Monitoring6
Thread-Adaptive: High-Throughput Parallel Architectures of SLH-DSA on GPUs6
Exploring the DIMM PIM Architecture for Accelerating Time Series Analysis6
StreamDQ: HBM-Integrated On-the-Fly DeQuantization via Memory Load for Large Language Models6
High-Bandwidth Flash for KV Caches: Endurance and Performance Implications6
QuArch: A Question-Answering Dataset for AI Agents in Computer Architecture6
Enabling Computation and Communication Overlap in PIMs for On-Device LLM Inference6
Exploiting Intel Advanced Matrix Extensions (AMX) for Large Language Model Inference6
Improving Performance on Tiered Memory With Semantic Data Placement6
Accelerating Deep Reinforcement Learning via Phase-Level Parallelism for Robotics Applications6
NoHammer: Preventing Row Hammer With Last-Level Cache Management5
Mitigating Timing-Based NoC Side-Channel Attacks With LLC Remapping5
Hardware-Accelerated Parallel Wrong-Path Execution for Spectre Gadget Detection5
PUDTune: Multi-Level Charging for High-Precision Calibration in Processing-Using-DRAM5
Nighthawk: Zero-Copy Cache Quarantine for Invisible Speculation4
pNet-gem5: Full-System Simulation With High-Performance Networking Enabled by Parallel Network Packet Processing4
SparseLeakyNets: Classification Prediction Attack Over Sparsity-Aware Embedded Neural Networks Using Timing Side-Channel Information4
LADIO: Leakage-Aware Direct I/O for I/O-Intensive Workloads4
DeMM: A Decoupled Matrix Multiplication Engine Supporting Relaxed Structured Sparsity4
ReplayOpt: Optimizer-State Replay to Resolve Critical-Path Bottlenecks in Offloaded Training4
RAESC: A Reconfigurable AES Countermeasure Architecture for RISC-V With Enhanced Power Side-Channel Resilience4
High-Performance Winograd Based Accelerator Architecture for Convolutional Neural Network4
Efficient Deadlock Avoidance by Considering Stalling, Message Dependencies, and Topology4
Xami : E x pert-Aware A daptive Compression for Mi 3
ZoneBuffer: An Efficient Buffer Management Scheme for ZNS SSDs3
Hisui: Unlocking Tiered Memory Efficiency for FaaS Workloads3
SSD Offloading for LLM Mixture-of-Experts Weights Considered Harmful in Energy Efficiency3
Camulator: A Lightweight and Extensible Trace-Driven Cache Simulator for Embedded Multicore SoCs3
Exploring Volatile FPGAs Potential for Accelerating Energy-Harvesting IoT Applications3
A Quantum Computer Trusted Execution Environment3
Guard Cache: Creating Noisy Side-Channels3
Enabling Cost-Efficient LLM Inference on Mid-Tier GPUs With NMP DIMMs3
Adaptive Web Browsing on Mobile Heterogeneous Multi-cores3
Primate: A Framework to Automatically Generate Soft Processors for Network Applications3
KiF: Accelerating Low-Batch LLM Inference Using In-Flash KV Cache3
Memory-Centric MCM-GPU Architecture3
FPGA-Accelerated Data Preprocessing for Personalized Recommendation Systems3
Energy-Efficient Bayesian Inference Using Bitstream Computing3
Fast Performance Prediction for Efficient Distributed DNN Training3
LeakDiT: Diffusion Transformers for Trace-Augmented Side-Channel Analysis3
Driving the Core Frontend With LiteBTB3
A Flexible Hybrid Interconnection Design for High-Performance and Energy-Efficient Chiplet-Based Systems3
Enhancing the Reach and Reliability of Quantum Annealers by Pruning Longer Chains3
Agentic LLMs for Microarchitecture Research: The Role of the Human Architect3
H 3 : H ybrid Architecture Using H igh Bandwidth Memory3
Understanding the Performance Behaviors of End-to-End Protein Design Pipelines on GPUs3
SoftmaxPIM: An HBM-Based PIM Architecture for Accelerating GeMV–Softmax Execution Pipeline3
NDPool: Correctness-Preserving Shared Execution for Efficient LLM Inference on CXL-NDP Systems3
Fast Inter-Enclave Communication Encryption3
Per-Row Activation Counting on Real Hardware: Demystifying Performance Overheads2
LWAL: Lightweight Adaptive Learning-Driven Cache Bypassing for GPUs2
CABANA : Cluster-Aware Query Batching for Accelerating Billion-Scale ANNS With Intel AMX2
Enhancing DNN Training Efficiency Via Dynamic Asymmetric Architecture2
Redundant Array of Independent Memory Devices2
Analyzing and Exploiting Memory Hierarchy Parallelism With MLP Stacks2
Accelerating LLM Inference via Dynamic KV Cache Placement in Heterogeneous Memory System2
Computational CXL-Memory Solution for Accelerating Memory-Intensive Applications2
3DLS: A 3D Logic-Stacked Architecture for Disaggregated LLM Serving2
Accelerating Page Migrations in Operating Systems With Intel DSA2
IntervalSim++: Enhanced Interval Simulation for Unbalanced Processor Designs2
Reducing the Silicon Area Overhead of Counter-Based Rowhammer Mitigations2
Direct-Coding DNA With Multilevel Parallelism2
T-CAT: Dynamic Cache Allocation for Tiered Memory Systems With Memory Interleaving2
EgDiff: An Enhanced Global Load Value Predictor2
Approximate Multiplier Design With LFSR-Based Stochastic Sequence Generators for Edge AI2
Amethyst: Reducing Data Center Emissions With Dynamic Autotuning and VM Management2
MOST: Memory Oversubscription-Aware Scheduling for Tensor Migration on GPU Unified Storage2
On the Feasibility Boundaries of LLM Prefill Inference on Unified-Memory Edge SoCs2
Hungarian Qubit Assignment for Optimized Mapping of Quantum Circuits on Multi-Core Architectures2
Exploring High-Bandwidth Flash for Modern LLM Inference: Opportunities and Challenges2
PINSim: A Processing In- and Near-Sensor Simulator to Model Intelligent Vision Sensors2
Near-HBM Tensor Core Acceleration for Fine-Grained Sparse Matrix-Matrix Multiplication2
Architectural Implications of GNN Aggregation Programming Abstractions2
gem5-accel: A Pre-RTL Simulation Toolchain for Accelerator Architecture Validation2
R.I.P. Geomean Speedup Use Equal-Work (Or Equal-Time) Harmonic Mean Speedup Instead2
Cost-Effective Extension of DRAM-PIM for Group-Wise LLM Quantization2
Enhancing DCIM Efficiency with Multi-Storage-Row Architecture for Edge AI Workloads2
FullPack: Full Vector Utilization for Sub-Byte Quantized Matrix-Vector Multiplication on General Purpose CPUs2
On Internally Tagged Instruction Set Architectures2
Minimal Counters, Maximum Insight: Simplifying System Performance With HPC Clusters for Optimized Monitoring2
CGR-NPU: A Hybrid CGRA and NPU Architecture for Adaptive Neural Computing Workloads1
Efficient MoE Model Fine-Tuning on Commodity GPU Server With Offloading1
MajorK: Majority Based kmer Matching in Commodity DRAM1
X-PPR: Post Package Repair for CXL Memory1
Fusing Adds and Shifts for Efficient Dot Products1
A Case for Hardware Memoization in Server CPUs1
Clover: Storage-Efficient Page Table Replication in Wafer-Scale GPUs1
Tulip: Turn-Free Low-Power Network-on-Chip1
A Data Prefetcher-Based 1000-Core RISC-V Processor for Efficient Processing of Graph Neural Networks1
Balancing Performance Against Cost and Sustainability in Multi-Chip-Module GPUs1
HINT: A Hardware Platform for Intra-Host NIC Traffic and SmartNIC Emulation1
Supporting a Virtual Vector Instruction Set on a Commercial Compute-in-SRAM Accelerator1
BlockPIM: Enabling Sparse LLM Inference on Dense PIM Architectures1
A Hardware-Friendly Tiled Singular-Value Decomposition-Based Matrix Multiplication for Transformer-Based Models1
An Intermediate Language for General Sparse Format Customization1
Unleashing the Potential of PIM: Accelerating Large Batched Inference of Transformer-Based Generative Models1
eDKM: An Efficient and Accurate Train-Time Weight Clustering for Large Language Models1
Privilege Level Dynamics and Their Impact on Conditional Branch Prediction Performance in FaaS1
Architectural Security Regulation1
SPAM: Streamlined Prefetcher-Aware Multi-Threaded Cache Covert-Channel Attack1
GPU-Centric Memory Tiering for LLM Serving With NVIDIA Grace Hopper Superchip1
Contention-Aware GPU Thread Block Scheduler for Efficient GPU-SSD1
A Multiple-Aspect Optimal CNN Accelerator in Top1 Accuracy, Performance, and Power Efficiency1
Exploiting Intel AMX Power Gating1
InfAMAX: Bridging the Compute-Memory Gap in Intel AMX for Efficient LLM Inference1
Electra: Eliminating the Ineffectual Computations on Bitmap Compressed Matrices1
TeleVM: A Lightweight Virtual Machine for RISC-V Architecture1
Capacity-Latency Tradeoffs in CXL Memory Expander at Hyperscale1
MixDiT: Accelerating Image Diffusion Transformer Inference With Mixed-Precision MX Quantization1
Halis: A Hardware-Software Co-Designed Near-Cache Accelerator for Graph Pattern Mining1
Characterization and Analysis of the 3D Gaussian Splatting Rendering Pipeline1
Pyramid: Accelerating LLM Inference With Cross-Level Processing-in-Memory1
Address Scaling: Architectural Support for Fine-Grained Thread-Safe Metadata Management1
Heterogeneous Mapping for Analog In-Memory Computing Accelerators: A Unified Workflow1
Intelligent SSD Firmware for Zero-Overhead Journaling1
Cache and Near-Data Co-Design for Chiplets1
Dicemite: Scaling Shared-Memory Isolation for Highly Consolidated FaaS Workers1
Approximate SFQ-Based Computing Architecture Modeling With Device-Level Guidelines1
Exploiting Direct Memory Operands in GPU Instructions1
Low-Latency PIM Accelerator for Edge LLM Inference1
A Partial Tag–Data Decoupled Architecture for Last-Level Cache Optimization1
Canal: A Flexible Interconnect Generator for Coarse-Grained Reconfigurable Arrays1
GEMM the New Gem: The Inevitable Kernel and its Sensitivity to Compiler Optimizations and Libraries1
0.05794095993042