What Happens When You 'Lobotomize' an LLM?
Tomasz Kolinko’s Effort Engine ranks matrix multiplications by impact and skips low-value ones during inference—at 50% effort quality stays nearly indistinguishable from full inference; at 30% you get coherent answers with ~3× speedup. Failure modes differ from quantization, and heat maps show where outputs shift, demoed on a MacBook.