Qwen 3.8 Flash Next + HERMES AGENT = AWESOME LOCAL AI AGENTS!
Digital Spaceport · 17:36
Digitalspaceport demonstrates running Qwen 3.8 Flash Next (a newer, larger MoE-style model) on a quad RTX 3090 rig via vLLM, achieving roughly double the tokens/sec of Llama.cpp, and uses it in a Hermes agent setup to...