Skip to main navigation Skip to search Skip to main content

A 8.4-1023.2 TOPS/W 1-7b Reconfigurable Computing In-Memory Macro with Charge-sharing-based Weighted Accumulator

Research output: Journal Publications and ReviewsRGC 21 - Publication in refereed journalpeer-review

Abstract

SRAM-based analog computing-in-memory (ACIM) demonstrates outstanding efficiency. However, it faces three critical challenges: significant ADC overhead, high latency for multi-bit inputs, and limited read bitline voltage. To address these issues, this work proposes a multi-bit highly reconfigurable 256×128 in-memory computing array supporting 1-7b input, 2-4b weight, and 1-7b output. The design employs bit-slicing (BS) with a charge-sharing-based binary-weighted accumulator (BSCHA) and a reconfigurable (1-7b) in-memory analog-todigital converter (IMADC) with shared voltage references. Three key innovations are introduced: 1) The IMADC occupies only 3% area overhead, achieving a 9× improvement compared to previous IMADC; 2) The BSCHA reduces latency by 1.9× and 6.6× compared to traditional pulse-width modulation (PWM) and bit-slicing modes, respectively; 3) A dual-8T bitcell enabling ternary weight storage through a decoupled read path, integrated with a read wordline under-driven cascode technique, improves linearity of unit discharge current by 7× and increases the usable read bitline voltage by 3.5×. Using noise resilient training, we show software performance for a MLP on MNIST, VGG- 8 on CIFAR-10, and Vision Transformer (ViT) on CIFAR-100 with respective accuracy reductions of only 0.1%, 0.4% and 0.4% due to non-idealities. The proposed macro demonstrates high energy/area efficiency (1023.2 TOPS/W, 27 TOPS/mm2 at 1/2/1b and 8.4 TOPS/W, 0.014 TOPS/mm2 at 7/4/7b) in 65 nm CMOS. It increases throughput (by 1.9×) and linearity (by 23×) compared to input pulse-width modulation by using BSCHA. Compared to conventional BS with digital accumulation after ADC, this method has 1.5×/6.6× better normalized energyefficiency/ throughput by reducing ADC operations in macrolevel. The system-level hardware performance of VGG-8 on CIFAR-10 is evaluated through combined SPICE and NeuroSim simulations, demonstrating the significant 6× enhancements in normalized energy efficiency. 

© 2026 IEEE. All rights reserved, including rights for text and data mining and training of artificial intelligence and similar technologies. Personal use is permitted,but republication/redistribution requires IEEE permission. 
Original languageEnglish
Number of pages15
JournalIEEE Transactions on Circuits and Systems for Artificial Intelligence
DOIs
Publication statusOnline published - 18 May 2026

Funding

This work was sponsored in part by RGC (C7003-24Y) and Innovation technology Fund Mid-Stream Research program under Grant ITS/018/22MS.

Research Keywords

  • Charge-sharing
  • Computing In-memory
  • Dual 8T SRAM
  • In-memory ADC
  • Weighted accumulator

RGC Funding Information

  • RGC-funded

Fingerprint

Dive into the research topics of 'A 8.4-1023.2 TOPS/W 1-7b Reconfigurable Computing In-Memory Macro with Charge-sharing-based Weighted Accumulator'. Together they form a unique fingerprint.

Cite this