*Analog* Deep Learning Processors in 2010

From 2007 to 2011, my startup Lyric Semiconductor, Inc., funded by a combination of DARPA and venture capital, created the first (and so far the only) commercial analog deep learning processor. 440,000 analog transistors did the work that would have required ~30,000,000 digital transistors, providing ~10x better Joules/Ops power compared to a digital tensor processing core.

When we announced the technology, it captured people’s imaginations and received a great deal of press coverage across major technology news outlets including being featured as the top story in the technology sections of both the in the New York Times1 and Wall Street Journal2. It was also covered in Wired3, MIT Technology Review4, Scientific American, The Register5, EE Times6, Reuters7, PC World8, The Flash Memory Summit9 (major applications in early flash memory drives), Phys Org10, Chip Estimate11, and then was picked up and spread worldwide by many others such as KD Nuggets12 and The Bulletin13.

Our analog tensor processing core was designed as a “plug and play” IP block for use within our digital deep learning microchips. (See this post about our overall deep learning processor architecture.)

Figure: the abstract of our publication in ACM/IEEE.

We published in ACM14. We also patented early explorations at MIT15, and then later the analog computing unit architecture16, analog storage17, factor/tensor operations18, error reduction of analog processing circuits19, I/O20, application in a signal processing pipeline21, and research on stochastic spiking circuits22. We used the terms “factor” and “tensors” interchangeably23.

I was the Founder, Chief Scientist, and CEO of Lyric Semiconductor, Inc. which grew out of my PhD work at MIT and Mitsubishi Electric Research Labs (MERL) in Cambridge, MA. Our small team comprised a little over 20 people people including David Reynolds, Jeff Bernstein, Bill Bradley, and Theo Weber.

Our startup was acquired by ADI, the largest analog/mixed-signal semiconductor company in the US, and became the new machine learning/AI chip group. Our innovations were incorporated into wireless infrastructure, medical devices, self-driving cars, mobile devices, and other applications.

The main innovation involved taking advantage of the fact that weights and activations in deep learning models can be represented by (quantized into) 7-bit numbers without losing accuracy in the overall computation.

The noise we observed in our circuits, was about 128th of our 1.8V power supply, so we could replace an 8-wire digital bus with a single wire carrying analog current. In practice we used a differential pair – 2 wires – to represent our analog values more robustly. Today’s chips often keep an entire tensor of logits within a reasonable numerical scale, by have a single “exponent” for all of those logits. We accomplished that same thing by having all of the currents representing logits sum into a single tail current that we carefully controlled. We used this control not just to control the exponent for numerical stability, but also to compensate for variations in manufacturing process, voltage and temperature — PVT).

Still, dropping from eight wires to two wires may not seem like a big enough win to justify the effort of designing analog tensor processors? Why bother?

The real win (about 10x in ops per Joule) came from two things:

  1. You get to use fewer transistors in the multiply-and-add. The main kind of math that deep learning processors need to do is multiplication and addition. Instead of a few thousand transistors needed to multiply two 8-bit digital numbers, we could use just 6 transistors to multiply two analog numbers. 500x fewer!
  2. Less intense switching is low power. On average our analog wires were not switching between 0 and 1, they were varying between intermediate current values.

    In the digital version of the same processor that we built, switching our wires from 1.8V to 0V and back dissipated the majority of power in our processor. On any given digital wire, this happens about half of the times that the processor’s clock ticks.

    By contrast, in our analog version, because of the statistical distribution of weights and activations around VDD/2, on average the currents in our analog wires did not change as widely nor abruptly.

It’s quite likely that all of this would still work in a modern 1nm semiconductor chips.

In 2010, a small percentage of the world’s computing workloads involved deep learning. Today deep learning work loads are becoming a driver for global energy consumption. As digital techniques become highly optimized, the 10x efficiency available from analog processing could matter even more today than it did when we first built our chips.

  1. A Chip That Digests Data and Calculates the Odds, Ashlee Vance, New York Times, August 17, 2010. https://www.nytimes.com/2010/08/18/technology/18chip.html ↩︎
  2. Lyric’s Chips Probably Interest Spooks, Shoppers, Don Clark, Wall Street Journal, August 17, 2010. https://www.wsj.com/articles/BL-DGB-17239 ↩︎
  3. Probabilistic Chip Promises Better Flash Memory, Spam Filtering, Wired, August 17, 2010. ↩︎
  4. A New Kind of Microchip, Tom Simonite, MITTechnology Review, August 17, 2010. https://www.technologyreview.com/2010/08/17/90614/a-new-kind-of-microchip/ ↩︎
  5. DARPA Funds Mr Spock on a Chip, The Register, August 17, 2010. ↩︎
  6. Chip Odyssey 2006: The Start of the Modern Chip Era, Weckel, Alan, EE Times, October 29, 2024. ↩︎
  7. The Odds Are Good That Lyric Semiconductor Will Change Computing, Reuters, August 2010. ↩︎
  8. Chip Startup Developing Probability Processor, PCWorld (IDG Communications), August 18, 2010. ↩︎
  9. LDPC Error Correction Using Probability Processing Circuits, Vigoda, Benjamin, Flash Memory Summit, Session 201, August 19, 2010. ↩︎
  10. Computer Chip That Computes Probabilities and Not Logic, Phys.org, August 19, 2010. ↩︎
  11. MIT Spin-Out Lyric Semiconductor Launches a New Kind of Computing With Probability Processing Circuits, ChipEstimate, August 2010. ↩︎
  12. A Chip That Digests Data and Calculates the Odds, Vance, Ashlee, New York Times (via KDnuggets), August 17, 2010. ↩︎
  13. A Chip That Calculates the Odds, Vance, Ashlee, The Bulletin (via New York Times), August 18, 2010. ↩︎
  14. Low power logic for statistical inference, Vigoda, Benjamin and Reynolds, David and Bernstein, Jeffrey and Weber, Theophane and Bradley, Bill, Association for Computing Machinery, 2010. ↩︎
  15. Analog Continuous Time Statistical Processing, Vigoda, Benjamin and Gershenfeld, Neil, Massachusetts Institute of Technology, issued December 28, 2010. US Patent 7,860,687 B2. ↩︎
  16. Belief Propagation Processor, Reynolds, David and Vigoda, Benjamin, Mitsubishi Electric Research Laboratories / Analog Devices, Inc., issued August 5, 2014. US Patent 8,799,346 B2. ↩︎
  17. Storage Devices with Soft Processing, Vigoda, Benjamin and Bernstein, Jeffrey and Venuti, Jeffrey and Alexeyev, Alexander and Nestler, Eric and Reynolds, David and Bradley, William and Zlatkovic, Vladimir, Analog Devices, Inc., US Patent 9,036,420 B2, issued May 19, 2015. ↩︎
  18. Programmable Probability Processing, Bernstein, Jeffrey and Vigoda, Benjamin and Nanda, Kartik and Chaturvedi, Rishi and Hossack, David and Peet, William and Schweitzer, Andrew and Caputo, Timothy, Analog Devices, Inc., issued February 7, 2017. US Patent 9,563,851 B2. ↩︎
  19. Apparatus and Method for Reducing Errors in Analog Circuits While Processing Signals, Vigoda, Benjamin, Mitsubishi Electric Research Laboratories, Inc., issued August 31, 2010. US Patent 7,788,312 B2. ↩︎
  20. Signal Mapping, Vigoda, Benjamin and Bernstein, Jeffrey and Alexeyev, Alexander and Venuti, Jeffrey, Lyric Semiconductor, Inc. / Analog Devices, Inc., published November 4, 2010. US Patent Application 20100281089 A1 (granted as US8,572,144 B2). ↩︎
  21. Analog Signal Conversion, Alexeyev, Alexander and Bernstein, Jeffrey and Bradley, William and Vigoda, Benjamin and Weber, Théophane, Lyric Semiconductor, Inc. / Analog Devices, Inc., US Patent 8,344,924 B2, issued January 1, 2013. ↩︎
  22. Mixed Signal Stochastic Belief Propagation, Bernstein, Jeffrey and Vigoda, Benjamin and Reynolds, David and Alexeyev, Alexander and Bradley, William, Analog Devices, Inc., issued July 29, 2014. US Patent 8,792,602 B2. ↩︎
  23. In a factor graph computing belief propagation, we could have, for example, a softAND gate with incident edges A, B, C. Logically, C = AND(A,B), which yields the tensor or “factor” computation p_C = \sum_{A,B,C} \delta(C-AND(A,B)) p_A p_B.
    Accelerating Inference: towards a full Language, Compiler and Hardware stack, Hershey, Shawn and Bernstein, Jeffrey and Bradley, Bill and Schweitzer, Andrew and Stein, Noah and Weber, Théophane and Vigoda, Benjamin, NIPS Workshop on Probabilistic Programming, December 12, 2012. (arXiv:1212.2991) ↩︎