Really really fast GPIO on Linux

Introduction

I started down this latest rabbit hole when looking for a way to get high speed data from the "low speed" GPIO header of a Raspberry Pi. I wanted to see if I could drive a parallel Eink display directly from the RPI's GPIO header instead of using an external controller like the IT8951. This would allow for the creation of a very low cost, high speed Eink display adapter. Spoiler - my idea worked and I created a working board with the help of my friend Martin:


I've also used this direct GPIO idea to drive QSPI LCDs from the RPI header (RPI doesn't have native QSPI support):



Back to GPIO...

The "digital pins" of most MCUs/SoCs are usually controlled from a few memory mapped ports. These exist in a similar form on simple Microcontrollers like the AVR 8-bit series (e.g. Arduino UNO) and on more complex SoCs like the Allwinner A733 (e.g. OrangePi Zero 3W). There are groups of bits to control each pin's function (e.g. simple input/output or special functions like I2C, SPI, etc) and additional bits to enable features such as pull-up resistors or interrupts. For anyone who is familiar with the digitalWrite() and digitalRead() functions of Arduino, it might surprise you to know that you can read and write all the pins of a GPIO port in parallel with a single instruction. The old Arduino UNO LCD shields depended on this feature - they're connected via an 8-bit parallel bus and make use of the fact that you can write 8-bits at once to the AVR's GPIO ports. If not for this ability to bit-bang 8-bits at a time in software, updating the contents of the LCD would be incredibly slow if done with individual calls to digitalWrite().

Problem 1 - When the abstraction gets in your way

On bare metal devices like Arduino, there is no such thing as privilege level - your code is the only thing running on the device and can do whatever it wants with the hardware. Direct GPIO manipulation doesn't have any restrictions, so you can use it as you please. On multi-user operation systems like Linux, each user's activities are protected from affecting other users and from accessing system resources directly. This provides security and stability. Hardware is shared by all users, so device tree overlays (aka drivers) are the gatekeepers managing these resources. The standard Linux GPIO driver is GPIOD. It allows for using multiple pins in parallel, but in a very limited way. Other drivers help manage the 'special functions' of GPIO pins too (e.g. I2C, SPI, PWM). These protocols have additional hardware to implement them, but they can also be implemented purely in software if needed since a GPIO pin is a GPIO pin. Implementing one of these protocols in software is normally called bit-banging. The real challenge comes when your project needs to do something that isn't easy to accomplish with the existing drivers. Let's examine one such situation:

I like to use the Pimoroni Display HAT mini on my Linux SBCs. The inclusion of a QWIIC I2C connector and a few pushbuttons makes it really convenient to use for a lot of different projects. The only problem is that the LCD connections conflict with the standard "Linux way" of doing things (see the chart below taken from the pinout.xyz website):

On the RPI 40-pin GPIO header, pins 19/21/23/24/26 are normally used for SPI. On this particular board, the LCD's Data/Command (DC is connected to pin 21) line is using what is normally the SPI MISO pin. This is a problem. If your program enables the SPI driver, it will take control of pins 19/21/23 and not let you use any of them as GPIO input/output. There are parameters that can be passed to the SPI driver in the /boot/config.txt file to tell it to not use the MISO pin, but those parameters vary by implementation or may not exist. This is now a dead-end for many users - without the 'magic words' to add to the boot configuration, the LCD will not work. To be fair, Pimoroni supplies software to solve this issue on RPIs, but for other Linux SBCs you're on your own.

Problem 2 - When the abstraction isn't fast enough

For my parallel Eink project I needed 8-bit parallel data to go as fast as possible. The Eink panels can normally handle data rates of 60-100MHz. Even if I could use the GPIOD driver to deliver this data, it would be much too slow for the task. The layers of code between a GPIOD API call and the GPIO pin changing are too thick for doing high speed operations. Speed is relative, but the best that you can get with that driver on a fast machine is around 200K changes per second. With direct GPIO register writes and efficient code, it is possible to get more than 50M changes per second (on the RPI Zero 2W). The RPI4B is able to get closer to 75M. Here's example code for the inner loop of writing parallel data to RPI's GPIO registers:

On the BCMxxx (RPI SoCs), the GPIO pins are accessed through "set" and "clear" registers. This is a common hardware feature of GPIO implementations that allows for atomic operations. A straight through register would require reading the contents, modifying the changed bits and writing back the new value. With set/clr registers, a single instruction can set or clear any number of bits without having to read the current state of the other bits. The Allwinner A733 happens to have both types of GPIO access (direct write and set/clr), while the RPIs just offer set/clear. For this particular loop (above), it could be faster if the RPI had a direct data write register instead of set/clear since integer instructions to mask/combine bits are much faster than the write instruction to a virtual hardware address.

How to access GPIO registers from user space?

Working with memory mapped hardware from user space on Linux requires a bit of extra effort compared to bare metal programming. Each user's code runs in a virtual address space and cannot directly access physical addresses. For example, the A733's GPIO registers start at physical address 0x2000000. If you write a program which tries to write to addresses in that range, it will likely cause an exception because nothing is mapped there in your user's instance. The key is to use the memory mapping API to map physical addresses into your current virtual space. Linux provides the mmap() function to do this mapping. The memory mapping driver lives in the device tree and, as is true for almost everything in Linux, is accessed with a file handle. Here's how to map the physical GPIO registers into our virtual address space:

You pass the physical starting address and the size of the area you want to map. The return value is a pointer to a block of addresses that will control the GPIO registers. Another important aspect to notice - your process needed root privilege to access /dev/mem; this is normally not given to user IDs, so you'll need to run as root (sudo).

It looks more complicated than it is...

Once you have a pointer to the GPIO registers, you can configure the pins as needed for your project. For my uses, I only need simple input/output capability. If you look at the Linux GPIO driver source code, you may be a bit intimidated by its complexity. It implements almost all of the features of each SoC and has to manage multiple users. To configure and use GPIO pins for your project, you probably need only a tiny amount of code. Here's my pinMode() function to set a GPIO pin to either input or output on the A733:



The A733 has 4 bits assigned per GPIO pin, the 16 pins in each group are controlled from 2 32-bit configuration registers. GPIO mode 0 is input and mode 1 is output. The details for these registers are found in the A733 User Manual. 


I map GPIO ports B through L into an 8-bit value. e.g. Port B3 would be pin number 0x13. The code above shows (1+(pin>>4)) because the first 0x80 bytes are special registers and the GPIO control registers start at 0x100. I created a PORT_REG structure to simplify (and homogenize) the code for different SoCs. For the A733, it looks like this:



The gap member vars are unused spaces in the 0x80 bytes per GPIO port (B-L).


Let's do something interesting beyond 'blinky'


With the info above, you can use your Linux SBC to blink an LED at greater than 50MHz! There may be some people satisfied with that accomplishment, but I think most of us would like to go further. A useful project for this knowledge would be to get that Pimoroni Display HAT mini working on the OrangePi Zero 3W, so let's do that. I've already written a SPI LCD display library (https://github.com/bitbank2/bb_spi_lcd). It has the init sequences for many many displays and all sorts of font and 2D graphics functions for drawing. Early in the development of the library, I also had the need for custom interface options, so I added the ability to provide your own reset and data writing functions. We can make use of this feature to write a little bit of code that's specific to our target hardware and get things working. For our purposes, we'll need to write pinMode(), digitalWrite() and a bit-bang SPI write function. The first version of the bit-bang SPI code will be 'naive' and use the digitalWrite() function in the inner loop. It looks like this:


Using my fast digitalWrite() function, this code can't go tremendously fast. To fill the 320x240 LCD will black takes 478ms. That's an equivalent SPI clock speed of about 2.5MHz. If we make the inner loop a bit more clever by directly manipulating the GPIO registers, we can make it go faster. Here's the fast version of that function:


This version generates output that's equivalent to about 6MHz. In this case, the CPU hardware is limiting the speed. Writing directly to the GPIO data register creates significant delay on the A733 hardware. I'm not sure if this is because of the memory management unit or another reason. I also discovered that the SET and CLR registers on the A733 don't work. On the H618 (OrangePi Zero 2W), I saw similar delays in writing to the GPIO registers. I'll experiment on more SBCs and see if the Broadcom SoCs are an outlier for having fast GPIO.

N.B.
I created this same code for the Raspberry Pi Zero 2W and the results ran much faster than on the A733. The simpler code (using the digitalWrite() function) got the equivalent of 33MHz SPI and the fast version got 45MHz. This indicates that the SET/CLR registers on RPI's SoC have much less latency than on both of the Allwinner SoCs even though the Allwinner CPUs execute instructions faster.

Final Thoughts

The GPIO hardware on Linux SBCs is capable of much more than I've touched on here. My aim with this article is to give you a starting point to explore more of what you can do on your hardware.  By using mmap() from user mode, you can treat Linux more like a bare metal programming environment and create projects that would otherwise require a much more complicated path through Linux kernel drivers. This enables doing projects like my example - running a 'difficult' LCD display on any SBC or driving high speed parallel data without extra hardware.

The source code for the example project above can be found here:



Comments

Popular posts from this blog

How much current do OLED displays use?

Surprise! ESP32-S3 has (a few) SIMD instructions

Fast SSD1306 OLED drawing with I2C bit banging