← all posts

What Can I Make This Do?

Hackable harnesses, local LLMs, reverse-engineering tools. The hacker mindset intensifies πŸ§‘β€πŸ’»

Let's see what can I make this do

If you give your mom a new gadget, like a phone, she'll probably ask, "What does this do?" That's a totally normal question.

If you give a new gadget to a hacker, then the question is:

What can I make this do?1

That's how hacker and inventor Pablos Holman described the mindset of a hacker in a conversation with Tim Ferriss.

People are taking apart software, connecting games that have no business communicating, modifying agent harnesses, and running large models on consumer hardware.

Some of it is useful. Some of it is deeply cursed, and the legality is questionable.

Interesting. 😎

Context engineering is dead. Long live harness engineering.

Well, no, it isn't. I just wanted to use that headline. It is more like a Pokémon evolving. 🀣

Hackemon: prompt engineering evolves into context engineering, then harness engineering

Prompt engineering asks: what should I tell the model?

Context engineering asks: what information should it have, and when? Retrieval, memory, tools, and the precious attention budget.2

Harness engineering asks: what kind of environment should it operate in? How does it use tools, recover from failures, and manage permissions and costs?

The interesting part is how much of that environment we can now modify:

  • Claude Code mods are TypeScript functions that customize prompts, interfaces, and built-in behavior through its plugin system.3
  • DeepSeek Harness makes models, tools, sessions, agent loops, and other components replaceable plugins.4
  • OpenCode exposes runtime hooks around model requests, tool execution, and session handling.5

Their extension systems differ, but each lets us work below the conversation layer. We can change what happens before a tool executes or trim its output afterward, whether or not the model remembers to ask.

That opens up a few experiments:

  • Context compression: trim bloated tool results and measure what gets lost.
  • Model routing: select models using task characteristics and measured outcomes.
  • Subagent tracking: record what each agent did, spent, and actually verified.
  • Better tools: replace enormous search outputs with precise queries and structured results.

There's a difference between giving a model a screwdriver and redesigning the workshop it operates in.

Gentlemen, we can rebuild it. We have the technology...

Unfortunately, we spent the entire budget on API tokens. πŸ’Έ

Take REA β€” Reverse Engineer Anything. It's a collection of agent-oriented workflows for investigating software behavior, native binaries, Electron applications, and more. Its goal is to help an agent gather evidence, explain how something works, and build a compatible implementation, not magically recover the original source code.6

Or Ghidra Headless MCP, which exposes disassembly, decompilation, cross-references, scripting, and other reverse-engineering operations to MCP-capable agents.7

The models themselves are getting better at this kind of work, too. OpenAI reported an 88.0% single-attempt result for GPT-6 Astra on its SRE-Bench evaluation, rising to 99.2% within four attempts. Those are vendor-reported benchmark results, not proof that an agent can flawlessly reverse engineer arbitrary software in the wild.8

And if you want a less serious demonstration, may I present the game-mashing goblins.

Nobody asked for Minecraft inside Skyrim

But somebody built it anyway.

  • SkyCraft combines Minecraft's blocks and combat with Skyrim's world, NPCs, and quests. Both games run together, communicating through shared memory.9
    A Minecraft character and inventory inside Skyrim's Riverwood village
  • Universal Modder gives coding agents skills, tools, and a shared knowledge base for modding PC games: inspect the engine, build a mod, generate assets through fal.ai, and test it in-game. Findings become field notes for the next agent.10
    Minecraft's Steve flying above Los Santos with an elytra
  • 2010 Rust Rewrite Mashup combines Modern Warfare 2, Skate 3, and Minecraft in a Rust runtime. Shoot through Minecraft blocks with MW2 guns, then press J to switch to a skateboard. Apparently, two games weren't enough.11
    A soldier using Skate 3 movement on a snowy Modern Warfare 2 map

Sir, why do you have an assault rifle and skater moves? This is Minecraft. 🀣

These are experimental integrations with rough edges. The common question behind them: "What can I make this do?"

Goblin Engineering at its finest. 😎

The hardware gets the same treatment:

Can I run a gigantic LLM on my gaming PC?

We have MoE at home

At GTC 2025, NVIDIA CEO Jensen Huang said reasoning and agentic AI needed 100 times more compute than expected a year earlier.12 Since then, his favorite answer in interviews has still been more compute. 😏

No, we have MoE at home

As you look at the shiny NVIDIA H200 racks displayed, your mom suddenly says "No, we have MoE at home" πŸ₯²

Mixture-of-experts (MoE) models activate only a fraction of their experts per token. The challenge is keeping the right weights within reach.

Colibri (Released Jul 19) – takes advantage of that by keeping always-needed weights in RAM and streaming routed experts from disk. Its C engines support several model families and can run without a GPU, though waiting on disk reads can make things slow.13

Strata (Released Sep 24) – shares Qwen3.8-Flash-Next across a gaming PC's GPU and CPU, keeping frequently used experts on the GPU and others in system RAM. Its documented setup starts at 12 GB of VRAM and 32 GB of RAM, depending on the model variant, with local APIs for apps and coding agents.14

A few models Colibri supports, with Qwen3.8 also running in Strata. B means billion parameters:

Model Parameters Active per token
Qwen3.6-35B-A3B 35B 3B
Qwen3.8-Flash-Next 125B* 6B
DeepSeek V4 Flash 284B 13B
GLM-5.2 744B ~40B

Qwen3.8's 125B figure excludes another 51B in n-gram embeddings and 4B for multi-token prediction. These are model parameter counts; memory requirements depend on weight compression, offloading, and context length.

As an ARPG player, I am obligated to do the pros & cons for this build:

Pros:

  • Run supported models larger than your GPU's memory.
  • No per-token API bill for local inference.
  • No 5-hour usage limits; run it 24/7 if you want.
  • Prompts can stay on your machine.
  • Tinker with the model, engine, and hardware setup.

Cons:

  • You still pay for hardware and electricity.
  • It can be slow, sometimes 1–2 tokens/s.
  • Experimental: not all models are supported, and if you want to try others, you're on your own.

With great power comes great responsibility

Uncle Ben was right: more control also means more responsibility.

Code that intercepts prompts and tool calls can access sensitive data or alter actions. Anthropic explicitly warns that Claude Code mods have the same machine access as Claude Code and are not sandboxed.3

Reverse engineering and modding also have authorization, licensing, security, and maintenance considerations. Something being technically possible doesn't mean it's appropriate to deploy or redistribute.

I want more hackable systems but also clear permission boundaries, observable execution, tests, checkpoints, and rollback paths.

We just unlocked the potential

Hackers have always inspected systems, questioned their limits, and tried ridiculous things. AI agents can make the next experiment easier to attempt.

Hacking isn't the new way. Hacking is the old way.

We just unlocked the potential. 🀝


Footnotes

  1. Pablos Holman, interview on The Tim Ferriss Show, episode 827; full transcript, September 2025. ↩

  2. Anthropic, Effective context engineering for AI agents, September 29, 2025. ↩

  3. Anthropic, Customize Claude Code with mods, October 1, 2026. ↩ ↩2

  4. DeepSeek, DeepSeek Harness developer preview and source repository. ↩

  5. OpenCode, Plugin API and runtime hooks. ↩

  6. morluto, REA β€” Reverse Engineer Anything, source repository and documentation. ↩

  7. mrphrazer, Ghidra Headless MCP, source repository. ↩

  8. OpenAI, GPT-6 Astra: A new generation of intelligence, September 2026, SRE-Bench results. ↩

  9. chasmlol, SkyCraft, source repository. Gameplay screenshot from the project. ↩

  10. rehan-remade, Universal Modder, source repository and README, including the modding workflow and shared knowledge base. Image extracted from the project's gameplay demo. ↩

  11. chasmlol, 2010 Rust Rewrite Mashup, source repository and README, including the Minecraft map and Skate 3 mode. Demo screenshot credited to chasm, reproduced in Sovereign Magazine. ↩

  12. Jensen Huang, GTC 2025 Keynote, March 2025, around 12:55–13:09 in NVIDIA's transcript. The comparison concerns compute requirements for reasoning and agentic AI relative to expectations a year earlier. ↩

  13. JustVugg, Colibri, source repository and README, including its architecture overview and hardware-dependent measurements. ↩

  14. Niko1221, Strata and How does Strata work?, source repository, setup requirements, and architecture documentation. ↩

πŸ€–Ask my AI