A Custom Zig Inference Engine with Infinite Streaming Context & Native Memory | Charles Sullivan
Columbus AI Tinkerers · 38:52
This talk is a live demo where an independent developer shows off a custom-built local LLM inference engine (written in Zig, not using llama.cpp) that can be interrupted mid-generation without losing state, streams ar...