We make games.
Small, systemic games built in Unity and C#. Our own titles are in development; announcements land here first.
- Unity & C#, start to finish
- In-house cinematic camera tools
- Played by our own bots before people
- Built for clean launches
Miaw Works is an independent game studio and tools lab from Pune, India. We make games in Unity, and we build Miaw QA Lab: a seeded bot plays your build, detectors watch for trouble, and an AI that has to cite its evidence turns thousands of log lines into a short, ranked bug list.
navmesh_explorer follows. Magenta tile, hole, broken door, slow zone: four seeded bugs to find.This is a real QA Lab run folder: a 55-second sandbox playtest at seed 42. Run the triage. Lines that share a root cause become one report, ranked by a formula you can read, not by a model's mood.
qalab triage run samples/sample_run --provider nonereal output, no model needed. The sample is hand-made test data; benchmark numbers get published once they're measured.
QA Lab calls AI APIs at four points: writing reports, finding the right design doc, reading screenshots, and embedding logs for grouping. Every answer passes through code that checks it before anyone sees it. Click a stage.
// cluster QAL-0001, built by code, not the model signature: KeyNotFoundException @ SeededEnemyRegistry.Get evidence: seq 019, 026 actions: seq 004 move_to · 007 interact Door_02 · 017 move_to docs: sandbox_design.md#enemy-registry priority: P2 ← computed, read-only
{
"summary": "Enemy lookup throws for ids that were never registered",
"suspected_cause": "Spawner skips registration when a spawn point is missing",
"evidence": [19, 26],
"steps": [{"action_ref": 17, "text": "Walk to the spawner"}],
"severity": "S2"
}The model writes the words. It never picks the priority, and if its severity drifts far from the computed priority the report is flagged for review. Answers are cached in SQLite, so a re-run costs nothing.
The design doc says how a feature is supposed to work, so the report can say "expected vs actual" with a citation instead of guessing. Bar widths are illustrative; real retrieval scores come from qalab eval triage (hit@3).
Pixel checks are free and run on every frame. The vision model only gets frames where pixels can't tell, like text spilling out of a box, so API calls stay low. The eval reports VLM calls, latency and cost per run.
evidence 19found in events.jsonlevidence 26found in events.jsonlevidence 99not in this run → dropped, report flaggedaction 17bot really did move_to (20, 0, 9)"step: open the cellar door"no matching bot action → marked inferreddoc #enemy-registryretrieved for this clusterThis is the part that makes AI usable in QA: a hallucinated id can't reach a ticket.
Small, systemic games built in Unity and C#. Our own titles are in development; announcements land here first.
Miaw QA Lab: open-source, AI-assisted test automation for Unity. Records playtests, plays the build, and writes the bug reports.
We're looking for Unity teams to pilot QA Lab on a real project. We write the game adapter with you; your data can stay on your machines.
We also build in-house Unity tools for shooting cinematic shots: trailers, cutscenes and screenshots without hand-keying every camera. Try the shot types below, sketched in your browser.
Browser sketch of the shot types, not the Unity tool itself. Want to see the real thing? Ask for a reel.
Build the player if needed, run a bot playtest and a menu crawl, check the screenshots, triage both runs, open the report.
scripts\run_pipeline.ps1 -Seed 42 -Duration 120 -Open
A C# package writes every log, action, metric and screenshot to a run folder, thread-safe.
run.json · events.jsonlA seeded bot walks the NavMesh toward places it hasn't been, or clicks through your menus.
-qalabAdapterDetectors flag falls out of the world, getting stuck, tunnelling, frame spikes and exception bursts, each with a screenshot.
DetectorHubPixel checks catch magenta textures and black screens. A vision model looks only where they're blind.
qalab vision analyzeLogs are normalised, grouped by signature and ranked P1–P4. An LLM drafts each report; every id it cites is checked.
qalab triage runA sandbox with 16 seeded bugs of known cause lets us score grouping and detection against ground truth.
qalab evalA model that invents repro steps or priorities makes QA worse. These rules are enforced in code and tests, not in a slide deck.
weight × (1 + log2 count) × runs × crash decides P1–P4. The model may suggest severity; if it disagrees, the report is flagged.
Evidence, action and doc ids are checked against the run. Unknown ids are dropped; a repro step with no matching bot action is marked inferred.
--provider none writes template reports and retrieval falls back to TF-IDF. The model is an upgrade, not a dependency.
A detector that throws is switched off for the run. The game keeps playing and the run keeps recording.
Run a local model through Ollama and nothing leaves the machine. API keys live in .env, never in logs, caches or reports.
Every figure we publish comes from a script in the repo, with the command that reproduces it.
Miaw QA Lab is public on GitHub: the code, the specs, the data contracts, every design decision and every known limitation. Studios and reviewers can audit exactly how it works before talking to us.
git clone https://github.com/sorinnha/miaw-qa-lab
README.mdstart here: what it does, quick start, limitations
docs/ARCHITECTURE.mdhow the Unity and Python halves fit
docs/DECISIONS.md30+ decision records with rejected alternatives
docs/specs/contracts, Unity package, triage, vision, CI
schemas/the JSON data contracts both sides test against
samples/sample_run/the run folder the triage demo above uses
python/src/qalab/triage, RAG, LLM providers, vision, eval
unity/com.miawworks.qalab/the UPM package: recorder, bot, detectors
docs/GAME_INTEGRATION.mdadd it to your own game
docs/USER_GUIDE.mdfor QC testers reading the reports
.github/workflows/CI: ruff, pytest on Windows + Ubuntu, C# tests
LICENSEMIT, use it commercially
What's built, what's waiting on its first real Unity run, and what's next. Straight from the project plan.
| version | status | ships |
|---|---|---|
| v0.1.0 | built · tested | Unity run recording, log triage with AI reports and RAG over design docs and code |
| v0.2.0 | first Unity run pending | Seeded bot (NavMesh explorer, UI crawler), detectors, screenshots, JUnit results, one-command pipeline |
| v0.3.0 | built · not yet measured | Triage evaluation on a 20-seed benchmark: clustering precision, recall, F1 |
| v0.4.0 | built · not yet measured | Vision: pixel heuristics, a vision-language model and a hybrid policy, scored per label |
| v1.0.0 | next | Running in a real game through a game adapter, triaged runs, scanner report, demo video |
I make games, and I got tired of reading the same exception four hundred times. So I'm building the tool I wanted: one that plays the build, groups the noise, and never pretends it found more than it did.
Miaw Works is founder-led and built in the open. Every design decision behind QA Lab, with the alternatives we turned down, is written up in the repo, and so is every limitation.
Pilot QA Lab on your project. We'll write the game adapter together.
Game dev, tools, QA automation, applied ML. Say hello.
Fact sheet, logos and colours on the press page.