If you give your mom a new gadget, like a phone, she'll probably ask, "What does this do?" That's a totally normal question.
If you give a new gadget to a hacker, then the question is:
What can I make this do?1
That's how hacker and inventor Pablos Holman described the mindset of a hacker in a conversation with Tim Ferriss.
People are taking apart software, connecting games that have no business communicating, modifying agent harnesses, and running large models on consumer hardware.
Some of it is useful. Some of it is deeply cursed, and the legality is questionable.
Interesting. π
Context engineering is dead. Long live harness engineering.
Well, no, it isn't. I just wanted to use that headline. It is more like a PokΓ©mon evolving. π€£
Prompt engineering asks: what should I tell the model?
Context engineering asks: what information should it have, and when? Retrieval, memory, tools, and the precious attention budget.2
Harness engineering asks: what kind of environment should it operate in? How does it use tools, recover from failures, and manage permissions and costs?
The interesting part is how much of that environment we can now modify:
- Claude Code mods are TypeScript functions that customize prompts, interfaces, and built-in behavior through its plugin system.3
- DeepSeek Harness makes models, tools, sessions, agent loops, and other components replaceable plugins.4
- OpenCode exposes runtime hooks around model requests, tool execution, and session handling.5
Their extension systems differ, but each lets us work below the conversation layer. We can change what happens before a tool executes or trim its output afterward, whether or not the model remembers to ask.
That opens up a few experiments:
- Context compression: trim bloated tool results and measure what gets lost.
- Model routing: select models using task characteristics and measured outcomes.
- Subagent tracking: record what each agent did, spent, and actually verified.
- Better tools: replace enormous search outputs with precise queries and structured results.
There's a difference between giving a model a screwdriver and redesigning the workshop it operates in.
Gentlemen, we can rebuild it. We have the technology...
Unfortunately, we spent the entire budget on API tokens. πΈ
Take REA β Reverse Engineer Anything. It's a collection of agent-oriented workflows for investigating software behavior, native binaries, Electron applications, and more. Its goal is to help an agent gather evidence, explain how something works, and build a compatible implementation, not magically recover the original source code.6
Or Ghidra Headless MCP, which exposes disassembly, decompilation, cross-references, scripting, and other reverse-engineering operations to MCP-capable agents.7
The models themselves are getting better at this kind of work, too. OpenAI reported an 88.0% single-attempt result for GPT-6 Astra on its SRE-Bench evaluation, rising to 99.2% within four attempts. Those are vendor-reported benchmark results, not proof that an agent can flawlessly reverse engineer arbitrary software in the wild.8
And if you want a less serious demonstration, may I present the game-mashing goblins.
Nobody asked for Minecraft inside Skyrim
But somebody built it anyway.
- Universal Modder gives coding agents skills, tools, and a shared knowledge base for modding PC games: inspect the engine, build a mod, generate assets through fal.ai, and test it in-game. Findings become field notes for the next agent.10

- 2010 Rust Rewrite Mashup combines Modern Warfare 2, Skate 3, and Minecraft in a Rust runtime. Shoot through Minecraft blocks with MW2 guns, then press J to switch to a skateboard. Apparently, two games weren't enough.11

Sir, why do you have an assault rifle and skater moves? This is Minecraft. π€£
These are experimental integrations with rough edges. The common question behind them: "What can I make this do?"
Goblin Engineering at its finest. π
The hardware gets the same treatment:
Can I run a gigantic LLM on my gaming PC?
We have MoE at home
At GTC 2025, NVIDIA CEO Jensen Huang said reasoning and agentic AI needed 100 times more compute than expected a year earlier.12 Since then, his favorite answer in interviews has still been more compute. π
As you look at the shiny NVIDIA H200 racks displayed, your mom suddenly says "No, we have MoE at home" π₯²
Mixture-of-experts (MoE) models activate only a fraction of their experts per token. The challenge is keeping the right weights within reach.
Colibri (Released Jul 19) β takes advantage of that by keeping always-needed weights in RAM and streaming routed experts from disk. Its C engines support several model families and can run without a GPU, though waiting on disk reads can make things slow.13
Strata (Released Sep 24) β shares Qwen3.8-Flash-Next across a gaming PC's GPU and CPU, keeping frequently used experts on the GPU and others in system RAM. Its documented setup starts at 12 GB of VRAM and 32 GB of RAM, depending on the model variant, with local APIs for apps and coding agents.14
A few models Colibri supports, with Qwen3.8 also running in Strata. B means billion parameters:
| Model | Parameters | Active per token |
|---|---|---|
| Qwen3.6-35B-A3B | 35B | 3B |
| Qwen3.8-Flash-Next | 125B* | 6B |
| DeepSeek V4 Flash | 284B | 13B |
| GLM-5.2 | 744B | ~40B |
Qwen3.8's 125B figure excludes another 51B in n-gram embeddings and 4B for multi-token prediction. These are model parameter counts; memory requirements depend on weight compression, offloading, and context length.
As an ARPG player, I am obligated to do the pros & cons for this build:
Pros:
- Run supported models larger than your GPU's memory.
- No per-token API bill for local inference.
- No 5-hour usage limits; run it 24/7 if you want.
- Prompts can stay on your machine.
- Tinker with the model, engine, and hardware setup.
Cons:
- You still pay for hardware and electricity.
- It can be slow, sometimes 1β2 tokens/s.
- Experimental: not all models are supported, and if you want to try others, you're on your own.
With great power comes great responsibility
Uncle Ben was right: more control also means more responsibility.
Code that intercepts prompts and tool calls can access sensitive data or alter actions. Anthropic explicitly warns that Claude Code mods have the same machine access as Claude Code and are not sandboxed.3
Reverse engineering and modding also have authorization, licensing, security, and maintenance considerations. Something being technically possible doesn't mean it's appropriate to deploy or redistribute.
I want more hackable systems but also clear permission boundaries, observable execution, tests, checkpoints, and rollback paths.
We just unlocked the potential
Hackers have always inspected systems, questioned their limits, and tried ridiculous things. AI agents can make the next experiment easier to attempt.
Hacking isn't the new way. Hacking is the old way.
We just unlocked the potential. π€
Footnotes
-
Pablos Holman, interview on The Tim Ferriss Show, episode 827; full transcript, September 2025. β©
-
Anthropic, Effective context engineering for AI agents, September 29, 2025. β©
-
Anthropic, Customize Claude Code with mods, October 1, 2026. β© β©2
-
DeepSeek, DeepSeek Harness developer preview and source repository. β©
-
OpenCode, Plugin API and runtime hooks. β©
-
morluto, REA β Reverse Engineer Anything, source repository and documentation. β©
-
mrphrazer, Ghidra Headless MCP, source repository. β©
-
OpenAI, GPT-6 Astra: A new generation of intelligence, September 2026, SRE-Bench results. β©
-
chasmlol, SkyCraft, source repository. Gameplay screenshot from the project. β©
-
rehan-remade, Universal Modder, source repository and README, including the modding workflow and shared knowledge base. Image extracted from the project's gameplay demo. β©
-
chasmlol, 2010 Rust Rewrite Mashup, source repository and README, including the Minecraft map and Skate 3 mode. Demo screenshot credited to chasm, reproduced in Sovereign Magazine. β©
-
Jensen Huang, GTC 2025 Keynote, March 2025, around 12:55β13:09 in NVIDIA's transcript. The comparison concerns compute requirements for reasoning and agentic AI relative to expectations a year earlier. β©
-
JustVugg, Colibri, source repository and README, including its architecture overview and hardware-dependent measurements. β©
-
Niko1221, Strata and How does Strata work?, source repository, setup requirements, and architecture documentation. β©


