Solving a puzzle with OpenCode - part 1
Background
Several years ago I put together a “puzzle” that would take the solver through a series of data manipulation tasks to eventually reveal a message. Why? I do not have an answer for that; let’s move on.
I have had a couple of attempts to solve the puzzle programmatically myself - once memorably on a flight before holiday mode had fully engaged - and seeing as I know all the answers it shouldn’t be that hard, right? Well, there a couple of steps that need a human eye… or do they?
OpenCode is a coding agent that provides access to hosted models that fit nicely into my OOO life’s LLM budget of £0.00 per month, and with it I shall:
- Attempt to solve the puzzle with minimal, open-ended prompting.
- Attempt to limp through trickier stages with targeted prompt plans.
I will be running the OpenCode coding agent on a Thinkpad x250 (i7-5600U / 8GiB / Intel graphics / mechanical HDD), running Ubuntu. Fairly modest specs for 2026 in other words.
If you want to try this yourself please be aware that the entry point URL to the puzzle itself does contain a flashing image. The flashing image does not feature in this article.
Let’s go
Straight out of the gate.
It did not take long for the Big Pickle model to get further than ChatGPT did when I attempted similar in 2025. The animated gif from the URL was extracted fairly quickly, and exploded into separate frames.
Each frame contained a QR code that decodes into an (almost) base64 string. I say almost as I had added a prefix to each base64 string to indicate the ordering of the frames, should the solver’s chosen explosion method not preserve it.
eg.
# frame 1
1:dGVzdCBzdGF0ZW1lbnQK...
# frame 2
2:YW5vdGhlciBzdGF0ZW1l...
With : being very much not a base64 character it
was hoped this would act as a clue to humans.
OpenCode was immediately able to strip these, concatenate the
contents of the frames - crucially in the correct order - and
decode. One technique to be relied on throughout the puzzle if
doing it by hand is using file to infer file types,
and inform subsequent steps. eg.
file decoded_file
decoded_file: XZ compressed data, checksum CRC64
This was of course second nature to the model, and there were no signs of tripping up here.
Decompressing the .xz file reveals a
.jpg, which is there things get interesting. This
was about as far as my solver script from the past had
managed.
When generating this image for the puzzle, I made sure to
pick characters and a font (Chilanka-Regular) that
I envisioned being awkward for OCR tool. Certainly, I was never
able to reverse engineer it satisfactorily with
tesseract in the past, and so was particularly keen
to see how an LLM would get on.
OpenCode appeared to get the memo however and attempted OCR across multiple fonts. This took some time, with some interesting commentary.
Hmm!!! 'a' (double-story) has a stem that ends at the baseline, sometimes with a small tail. Hmm! ✓
Hmm! 'Q' has a tail. Hmm! (cap height, excluded)
Hmm! 'o' has no spur. ✗
Hmm! Hmm, but the counter analysis says the counter is open at the bottom-right. Hmm! For a double-story 'a', the counter (bowl) is closed. Hmm! ✗
Hmm!!! ARGH.
Hmm, OK. Let me look at this from the pixel data ONE more time, very carefully, and just carefully describe g0:
14 ............#######....... cols 12-18
15 ..........###########..... cols 10-20
16 ........###############... cols 8-22
17 ......##################.. cols 6-23
18 .....#######.......#####.. cols 5-11, 19-23
19 ....######..........#####. cols 4-9, 20-24
20 ...######............####. cols 3-8, 21-24
21 ...#####.............####. cols 3-7, 21-24
22 ..#####...............#### cols 2-6, 22-25
23 ..####................#### cols 2-5, 22-25
Assumptions were made
There was some side-quest snooping around the site in general, which ended up influencing some decisions negatively later on.
Key find: the CV says "fan of regex". The answer is a regex. Let me read the full CV and test the string against the site.
113,414 tokens and 2hr 15min later
Wow! While the text was successfully extracted, the model did not attempt to reason about what its purpose would be beyond its regex assumption.
Either way I was impressed - excited even - when I read into
the process: ranking the likelihood of matches across >2,000
typefaces, on a glyph by glyph basis using a combination of
PIL/numpy, gocr and
tesseract!
Solved. Final transcription: aGt@!!?eeP4565tql
The chain
1. puzzle.html contains only <h1>Puzzle</h1> + puzzle.gif (no hint text).
2. puzzle.gif = 7-frame animated QR. zbarimg decoded 7 base64 payloads (2,500 chars ×6 + 2,232 = 17,232).
3. base64 → 12,924 bytes starting fd 37 7a 58 5a 00 = XZ; xz -t passes.
4. Decompress → 13,079-byte grayscale JPEG, 531×67 (answer.bin).
5. Threshold at 128 → exactly 17 glyphs, laid out with a constant 6px inter-glyph gap (so advance widths carry no information — the author re-spaced the glyphs).
Pressing on
Despite the confidence shown in the output above, the puzzle is not fully solved, however I did ask for a writeup README.
Of the people I have shared the exercise with in the past this is about as far as they have got too. While we did not arrive at the full solution through a one-shot prompt, I am encouraged by the results and am eager to see how far we can take this with the technology at our disposal.
In the next part we shall resume with a fresh context. In the meantime:
- How long does your coding agent take to get to this point?
- Can you solve the remainder of the puzzle?
All that is left now if for me to bestow the highest possible affirmation that engineer can upon OpenCode for its performance so far, shown below.