OpenTPU – An open-source AI accelerator, developed by AI
The repo, published by GitHub user FeSens, packs the full stack into one codebase: SystemVerilog RTL, a custom instruction set, a bit-exact Python simulator, a kernel compiler, a profiler called Lens, and host utilities including otpu-chat. The hardware target is a Xilinx Kintex-7 xc7k480t FPGA on an Inspur YPCB-00338 PCIe card with 4 GiB of DDR3-1066 memory across two channels. The design keeps things simple. A sequencer feeds a matrix unit that multiplies int8 weights streamed from DRAM, a vector unit handles fp32 math, and a quantizer converts results back to int8. There is no cache and no hidden scheduling, so every cycle can be traced. The project explicitly tests whether AI agents can design the chip that runs their own inference, extending the earlier auto-arch-tournament experiment. Commenters on Hacker News focused on whether AI-designed hardware is feasible and how it gets verified. For founders, the takeaway is that agents can now produce a working, readable accelerator stack on low-cost FPGA hardware. That makes custom inference hardware cheaper to prototype, although an FPGA demo is still far from a production chip.