About 43 million people worldwide are blind, many of them because the photoreceptors in their eyes have stopped working even though the rest of the visual pathway is intact. Retinal implants get around this by stimulating the surviving neurons directly, producing phosphenes, small points of light the brain reads as visual information. The catch is that standard frame-based image processing draws too much power and takes up too much space to fit inside an implantable device.
This project looks at a different approach: a delay-based architecture, where information is carried by the timing of signals instead of voltage levels, closer to how neurons actually communicate. We built an end-to-end pipeline on the NVIDIA Jetson Orin Nano that takes a video feed, down samples it, quantizes it down to a few brightness levels, encodes it into spikes for a spiking neural network, and drives a stimulation pattern for a retinal prosthesis. The down sampling step alone cuts the spatial data footprint by around 400x, and we can play it back at over 15 frames per second, fast enough for real-time, comfortable scene perception. The goal is a working baseline other researchers can build on for energy-efficient, delay-based bionic vision.
Mock event camera. We built software that mimics a real event-based (DVS) camera, taking standard video and converting it into sparse, quantized packets of information through down sampling, 4-level brightness and color quantization, and frame differencing. The output is a 96×54 spike tensor, close to what an actual DVS sensor produces.
Spiking neural network. That event data feeds into an SNN, which represents information as the timing of voltage pulses rather than floating-point numbers. That's both closer to how biological neurons work and far cheaper to compute, which is the whole point when you're trying to run this on something small enough to implant or wear.
Hardware. Most work in this space assumes a full workstation with a dedicated GPU, which isn't something a person can wear day to day. We built and tested the pipeline on an NVIDIA Jetson Orin Nano instead, since it's compact and efficient enough for local, real-time use.
A single 1920×1080 frame reduces to a 96×54 quantized matrix, a 400x cut in pixel count. Add 2-bit quantization and a one-minute clip that started at 40MB comes out to roughly 1-2MB, a 27x reduction in our current unpacked storage format. That's not the ceiling either: once we bit-pack the 2-bit data for the hardware version we're building next, the same file drops to around 390KB, closer to 100x.
On the SNN side, we implemented a leaky integrate-and-fire model in PyRTL and used it to build a neuro-inspired edge detection filter based on center-surround receptive fields. This was deployed to an Arty A7 100T FPGA, generating Verilog from PyRTL to test the SNN on real hardware.
We built a working pipeline that turns ordinary video into sparse, quantized spike data, cutting pixel count by 400x and file size by up to 27x in its current form. It shows this event-based approach can meaningfully shrink the data a bionic vision system needs to process in real time, which is exactly what you need for something low-power and wearable.
Next is bit-packed storage, this will enable us to bolster our 27x to use our DVS emulation pipeline directly to asynchronosuly train SNN, and validate the full software path. All this to impleent testbenches for timing jitter and scalability, to evaluate it's viability as a delay-based computing platform for real-time bionic vision.