IEEE Computer Architecture Letters

Papers
(The TQCC of IEEE Computer Architecture Letters is 3. The table below lists those papers that are above that threshold based on CrossRef citation counts [max. 250 papers]. The publications cover those that have been published in the past four years, i.e., from 2022-08-01 to 2026-08-01.)
ArticleCitations
Toward Practical 128-Bit General Purpose Microarchitectures112
Characterization and Analysis of Text-to-Image Diffusion Models37
Exploration of Algorithm-Hardware Co-Design for Floating-Point Digital Compute-in-Memory27
Old is Gold: Optimizing Single-Threaded Applications With ExGen-Malloc25
Accelerating Programmable Bootstrapping Targeting Contemporary GPU Microarchitecture23
A Characterization of Generative Recommendation Models: Study of Hierarchical Sequential Transduction Unit22
NeuroMTA: Programmable Simulation Framework for Multi-Tile NPU Architectures19
The Architectural Sustainability Indicator15
SCALES: SCALable and Area-Efficient Systolic Accelerator for Ternary Polynomial Multiplication13
Breaking the HBM Bit Cost Barrier: Domain-Specific ECC for AI Inference Infrastructure12
De-Quantization Penalties for Interactive LLM Inference on Prosumer GPUs12
Context-Aware Set Dueling for Dynamic Policy Arbitration12
MoSKA: Mixture of Shared KV Attention for Efficient Long-Sequence LLM Inference11
A Quantitative Analysis of Mamba-2-Based Large Language Model: Study of State Space Duality11
In-Depth Characterization of Machine Learning on an Optimized Multi-Party Computing Library10
SoCurity: A Design Approach for Enhancing SoC Security9
OASIS: Outlier-Aware KV Cache Clustering for Scaling LLM Inference in CXL Memory Systems9
Time Series Machine Learning Models for Precise SSD Access Latency Prediction9
Improving Energy-Efficiency of Capsule Networks on Modern GPUs9
Wafer-Scale GPU Memory Pool With In-Package Optics for Enhanced Capacity and Bandwidth8
AiDE: Attention-FFN Disaggregated Execution for Cost-Effective LLM Decoding on CXL-PNM8
Exploring KV Cache Quantization in Multimodal Large Language Model Inference8
In-Memory Versioning (IMV)7
REDIT: Redirection-Enabled Memory-Side Directory Architecture for CXL Memory Fabric7
RouteReplies: Alleviating Long Latency in Many-Chip-Module GPUs7
Straw: A Stress-Aware WL-Based Read Reclaim Technique for High-Density NAND Flash-Based SSDs7
A Case for In-Memory Random Scatter-Gather for Fast Graph Processing7
A Flexible Embedding-Aware Near Memory Processing Architecture for Recommendation System7
Improving Performance on Tiered Memory With Semantic Data Placement6
Accelerating Deep Reinforcement Learning via Phase-Level Parallelism for Robotics Applications6
Disaggregated Speculative Decoding for Carbon-Efficient LLM Serving6
Reducing Metadata and Page Migration Overheads in CXL-Based Secure Tiered Memory6
Security Helper Chiplets: A New Paradigm for Secure Hardware Monitoring6
Thread-Adaptive: High-Throughput Parallel Architectures of SLH-DSA on GPUs6
Exploring the DIMM PIM Architecture for Accelerating Time Series Analysis6
StreamDQ: HBM-Integrated On-the-Fly DeQuantization via Memory Load for Large Language Models6
High-Bandwidth Flash for KV Caches: Endurance and Performance Implications6
QuArch: A Question-Answering Dataset for AI Agents in Computer Architecture6
Enabling Computation and Communication Overlap in PIMs for On-Device LLM Inference6
Exploiting Intel Advanced Matrix Extensions (AMX) for Large Language Model Inference6
NoHammer: Preventing Row Hammer With Last-Level Cache Management5
Mitigating Timing-Based NoC Side-Channel Attacks With LLC Remapping5
Hardware-Accelerated Parallel Wrong-Path Execution for Spectre Gadget Detection5
PUDTune: Multi-Level Charging for High-Precision Calibration in Processing-Using-DRAM5
High-Performance Winograd Based Accelerator Architecture for Convolutional Neural Network4
Efficient Deadlock Avoidance by Considering Stalling, Message Dependencies, and Topology4
Nighthawk: Zero-Copy Cache Quarantine for Invisible Speculation4
pNet-gem5: Full-System Simulation With High-Performance Networking Enabled by Parallel Network Packet Processing4
SparseLeakyNets: Classification Prediction Attack Over Sparsity-Aware Embedded Neural Networks Using Timing Side-Channel Information4
LADIO: Leakage-Aware Direct I/O for I/O-Intensive Workloads4
DeMM: A Decoupled Matrix Multiplication Engine Supporting Relaxed Structured Sparsity4
ReplayOpt: Optimizer-State Replay to Resolve Critical-Path Bottlenecks in Offloaded Training4
RAESC: A Reconfigurable AES Countermeasure Architecture for RISC-V With Enhanced Power Side-Channel Resilience4
Understanding the Performance Behaviors of End-to-End Protein Design Pipelines on GPUs3
SoftmaxPIM: An HBM-Based PIM Architecture for Accelerating GeMV–Softmax Execution Pipeline3
NDPool: Correctness-Preserving Shared Execution for Efficient LLM Inference on CXL-NDP Systems3
Fast Inter-Enclave Communication Encryption3
Xami : E x pert-Aware A daptive Compression for Mi 3
ZoneBuffer: An Efficient Buffer Management Scheme for ZNS SSDs3
Hisui: Unlocking Tiered Memory Efficiency for FaaS Workloads3
SSD Offloading for LLM Mixture-of-Experts Weights Considered Harmful in Energy Efficiency3
Camulator: A Lightweight and Extensible Trace-Driven Cache Simulator for Embedded Multicore SoCs3
Exploring Volatile FPGAs Potential for Accelerating Energy-Harvesting IoT Applications3
A Quantum Computer Trusted Execution Environment3
Guard Cache: Creating Noisy Side-Channels3
Enabling Cost-Efficient LLM Inference on Mid-Tier GPUs With NMP DIMMs3
Adaptive Web Browsing on Mobile Heterogeneous Multi-cores3
Primate: A Framework to Automatically Generate Soft Processors for Network Applications3
KiF: Accelerating Low-Batch LLM Inference Using In-Flash KV Cache3
Memory-Centric MCM-GPU Architecture3
FPGA-Accelerated Data Preprocessing for Personalized Recommendation Systems3
Energy-Efficient Bayesian Inference Using Bitstream Computing3
Fast Performance Prediction for Efficient Distributed DNN Training3
LeakDiT: Diffusion Transformers for Trace-Augmented Side-Channel Analysis3
Driving the Core Frontend With LiteBTB3
A Flexible Hybrid Interconnection Design for High-Performance and Energy-Efficient Chiplet-Based Systems3
Enhancing the Reach and Reliability of Quantum Annealers by Pruning Longer Chains3
Agentic LLMs for Microarchitecture Research: The Role of the Human Architect3
H 3 : H ybrid Architecture Using H igh Bandwidth Memory3
0.18596315383911