Fourteen entries covering 21 repositories. Each is written from the repository itself, file tree, dependency files, migrations, rather than from memory. Commit counts are from the default branch. Where a repository is private there is no link, because there is nothing you could open.
2026-05 → 2026-06Sentinel-City
An agentic disaster-response platform where two reasoning loops detect, cordon and dispatch, and the operator sets the bounds rather than pressing the button.
- Repos
- Stack
- FastAPI · PostgreSQL 16 · LangGraph · Vertex AI Gemini · React 19 · Vite · Leaflet · Expo / React Native · Docker Compose
Architecture
Three clients over one system of record. A React 19 operator console on Leaflet, an Expo React Native app carrying three distinct roles, citizen, responder, admin, and a FastAPI backend holding disasters, citizen reports, field reports, emergency calls, stations, and simulated weather and traffic that react to whatever is currently active. Postgres is the single source of truth, migrated by numbered SQL files rather than an ORM migration tool, which keeps the schema legible as a document.
The agent work is a package, not a script. backend/pipeline/ separates the stages that a naive implementation would fuse into one prompt: extract pulls structure out of incoming reports, cluster groups them spatially, decide makes the call, dispatch_agent and execute carry it out, world_slice assembles the bounded view of the city that a model is allowed to reason over, and operator handles the human side. prank_check exists because a system that accepts citizen reports will receive false ones, and that filter is a separate, testable module rather than an instruction buried in a prompt.
Provenance is enforced at the write layer. Every agent-originated write is stamped with its source, and the mobile warning feed is server-filtered to agent-authored records only across five different tables, alerts, cordons, declared incidents, active dispatches, weather. Operator-drawn entries never reach citizens. That is a data-model decision rather than a UI decision, which is why it holds.
The bounding is explicit. backend/safety/policy.py is a standalone policy module with its own test file, test_policy_validators.py, sitting alongside eight further test modules covering extraction, clustering, decision, execution, geometry, world slicing, citizen reports and wiring, plus a separate stress test. The agents are the part that improvises; the policy layer is the part that does not.
The frontend does real work offline. A 6.4 MB road graph for the target city is baked at build time by a script and shipped as a static asset, so routing and avoidance run against a local graph instead of a round trip. A service worker caches map tiles. A 135 KB citizen simulation engine with its own spatial index drives the synthetic population.
Trade-offs
- Numbered SQL migrations instead of an ORM migration framework: harder to generate, trivial to read and to reason about in review. For a schema that four different clients depend on, being able to read the whole history as twelve short files is worth losing the generator.
- Baking the road graph into a static asset costs 6.4 MB of transfer and a build step, and buys routing that does not depend on an external service being up during an incident. For a disaster-response demo, the failure mode you are showing off must not itself depend on the network.
- Splitting the pipeline into nine modules is more code than a single orchestrating prompt and considerably more test surface. It is what makes the behaviour attributable, when a dispatch is wrong you can tell whether extraction, clustering or the decision produced it.
- Simulated weather and traffic that respond to active events add a whole subsystem that a real deployment would replace with feeds. Without it there is nothing to demonstrate, because real disasters cannot be scheduled.
What is not claimed hereThe repository README describes an earlier structure (orchestrator.py, agent_tools.py, api_client.py) that no longer exists. The description above follows the current file tree.
2024-11 → 2025-06Vigi, a command-line assistant, five times
Two files and 11 KB in November 2024. 322 KB and a project-generating agent by June 2025. Five repositories, each a restart rather than a branch.
- Repos
- Stack
- Python 3.9+ · LangGraph · Google Gemini · Groq · Typer · rich · questionary · Docker SDK
Architecture
The final generation is a CLI that dispatches to four substantially different subsystems behind one entry point. shell_smart is a LangGraph-driven interactive shell that generates commands from intent, explains them, validates them for safety, offers execution, summarises the output, and checks and installs missing dependencies with approval. developerch is a software-development agent that plans a project, generates its files, and can then be talked to about the codebase it produced. docker_part handles Docker in natural language, including interactive Dockerfile generation. chat_manage and convo_manage run persistent REPL sessions with history.
The persona and procedure system is the piece that generalises. Personas are JSON files on disk that the user can create; procedures are user-written Python functions in a known directory that the model is allowed to call as tools. Extension does not require touching the codebase, which is the same delegation instinct as the autograders, applied to the tool itself.
Approval is a structural boundary throughout. Commands are explained before they run, dependency installation asks first, and generated project state is written to a .vigi_dev_meta folder inside the target project so the agent's working memory lives next to the thing it is working on rather than in a global store.
The lineage is legible because each generation is a separate repository. Generation one is main.py plus temp_storage.py, 11 KB of Python total. Generation two is the first LangGraph rewrite. The final one is 322 KB across seventeen modules, with two different shell implementations kept side by side, shell_smart for the current one and shell_part for the older, still reachable under a different flag.
Trade-offs
- Restarting in a new repository rather than refactoring in place: the history of what changed between generations is lost, and in exchange every generation stays runnable and comparable. The five repositories are a record of a learning curve that a single rebased branch would have erased.
- Two shell implementations shipped in the same package is duplication, retained deliberately so the older behaviour stays available while the LangGraph one matures.
- Multi-provider support, Gemini as primary, Groq optionally for the Docker module, costs an abstraction the project did not strictly need at that size, and buys the ability to move when a provider degrades.
- Executing model-generated shell commands is the central risk of the whole idea. It is handled with explanation, validation and an interactive prompt rather than a sandbox, which is a reasonable trade for a personal tool and would not be one for a shared deployment.
2026-04 → 2026-05Pulse KPI
An HR performance platform that pulls tasks from whichever project tracker the company happens to use, scores them, and keeps the scoring code ignorant of where the tasks came from.
- Repos
- pulse-kpi, private, 82 commits
- Stack
- Django 4 · Django REST Framework · Django Channels · SimpleJWT · React 19 · Vite · React Router 7 · Google Gemini · Redis · PostgreSQL · Docker
Architecture
The central design decision is an adapter seam. ClickUp and a self-hosted Plane instance are both reduced to the same internal task shape by separate clients, and generate_kpi_for_member() never learns which one it is looking at. The source is selectable per request through settings, a CLI flag, or an API field, so switching a company from one tracker to another is configuration, not a rewrite. Local database tasks are a third source through the same seam.
Scoring is a model call over structured input, across eight evaluable categories and seven non-evaluable ones. Weights live in the database with their own migration (0006_kpiweights) rather than in code, which means the thing most likely to be argued about is the thing easiest to change.
Attendance is a rules engine, not a timestamp log. attendance_rules.py and attendance_rules_views.py sit alongside a recompute-attendance management command, meaning rules can change and history can be recomputed against them, which is the difference between a policy system and a spreadsheet.
Authentication is cookie-based JWT rather than tokens in local storage. Real-time updates run over Channels, with Redis in production and an in-memory layer in development so the local setup does not require Redis to be installed.
Nineteen migrations across the visible history, including a project-hierarchy and roles migration that arrives late, the shape of an org model that was discovered rather than designed up front.
Trade-offs
- A tracker-agnostic adapter is more code than integrating directly with one API, and it is the reason a second integration cost a client class instead of a fork.
- KPI weights in the database rather than in code: no type checking, no code review on a weight change, and no deploy needed either. For a number that HR will want to tune, that is the right side of the trade.
- SQLite in development with Postgres in production is a real correctness risk at the margins, taken in exchange for a zero-setup local environment.
- Letting a language model produce performance scores about people is the significant judgement call in this system. The structure around it, fixed categories, database-held weights, separate evaluable and non-evaluable groupings, an assignment and notification model for who evaluates whom, reads as an attempt to make the model one input to a reviewable process rather than the arbiter.
2025-11 → 2026-04Visual test builder
Drag nodes on a canvas, get a working Playwright script. A React Flow front end over a Django API, with the model providers behind a registry so none of them is load-bearing.
- Repos
- selenium-builder, private, 59 commits
- Stack
- Django 5 · Django REST Framework · React 19 · React Flow · Vite · Playwright (Python) · Docker Compose · GitHub Actions
Architecture
A monorepo with one job: let someone who does not write code assemble a browser test, and emit real Playwright Python from it. The canvas is React Flow; the backend owns the graph, the generated script, the test runs and their progress.
Model providers are a registry. ai/providers/ holds Anthropic, Gemini and OpenAI implementations behind a shared base class, resolved through registry.py. Nothing in the application imports a provider directly, so a provider change is a configuration change.
Generated commands are validated before they are trusted. ai/commands/schema.py defines the command vocabulary and ai/commands/validator.py, the largest module in that package, checks model output against it, with a pre_parser.py ahead of both. The model proposes a command; the schema decides whether it is a command at all. This is the same pattern as the retrieval work's confidence floor and Sentinel-City's policy module, arrived at independently in three different domains.
System prompts are a 13 KB Python module kept separate from the code that calls them, with a context.py that assembles per-request context. Request throttling is applied at the AI views specifically rather than globally.
Multi-tenancy arrives through migrations rather than a rewrite: user profiles, teams and team membership in 0005, invitations in 0006. Test runs get a progress field in 0003, a small migration that marks the point where runs stopped being instant.
A separate worker Dockerfile alongside the API image, so test execution scales independently of the web tier.
Trade-offs
- Generating Playwright Python rather than executing an interpreted graph directly: the output is a real file the user owns and can commit, run in their own CI, and edit by hand. The cost is that the generator has to produce code that is actually good, and every new capability needs both a node and a code template.
- A provider registry over three model vendors, at a project size where one would have done. It is cheap insurance and it makes provider comparison a runtime question.
- Schema-and-validator over trusting structured output: real work, and it is what stops a malformed generation from becoming a broken script the user has to debug without having written it.
- A separate worker image adds an operational component. Without it, a long test run blocks a web worker.
2026-05Range Alarm
A React Native alarm app that drops to a hand-written Kotlin module, because the one thing an alarm must do is the one thing JavaScript timers cannot.
- Repos
- Stack
- Expo SDK 54 · React Native 0.81 · TypeScript · Kotlin · expo-router · expo-sqlite · Zustand · Reanimated · d3-geo · TopoJSON · Luxon
Architecture
The project ejects to a bare Android workspace and ships a local native module, native-alarm, linked from the filesystem rather than a registry. 58 KB of Kotlin sits under a TypeScript app, a deliberate drop out of the cross-platform layer at exactly the point where the cross-platform layer cannot deliver. A JavaScript timer in a backgrounded React Native app does not survive Doze; an Android alarm scheduled through the platform does.
Time zones are treated as geography. world-atlas TopoJSON rendered through d3-geo and react-native-svg gives a real projected world map, with Luxon doing the zone arithmetic, so selecting a range is a spatial act rather than a dropdown.
State is Zustand; persistence is expo-sqlite on the device. There is no backend, which means no account, no sync, no server to keep running for a utility that must work on a plane.
Reanimated with the worklets runtime for interaction, expo-haptics for physical feedback, expo-audio for the alarm itself.
Trade-offs
- Writing and maintaining a Kotlin module inside an Expo project gives up the managed workflow, no Expo Go, a real Android build required. For an alarm, correctness at the moment of firing is the entire product, so the trade is not close.
- Local SQLite rather than a synced backend: no cross-device continuity, and no infrastructure, no account, and no privacy surface.
- Shipping a full world TopoJSON adds meaningful bundle weight in exchange for a map that works with no network, appropriate for an app whose main use case is travel.
- iOS is not addressed by the native module. The app is Android-first by construction.
2025-08 → 2025-11Carpet trade management
Stock, orders, lendings, contractors and payments for a physical business, with the money logic in a service layer instead of the route handlers.
- Repos
- Stack
- Python · Flask · JavaScript · GitHub Actions
Architecture
A strict two-layer split. app/api/ holds seven thin route modules, contractors, lendings, orders, payments, stock, stock reports, stock transactions, and app/services/ holds the logic they call. The route files are between 0.8 and 5 KB; order_service.py alone is 21 KB. Nothing that decides anything about money lives in a request handler.
Stock is modelled as transactions with a separate reporting service over the top, rather than as a mutable quantity column. That is the difference between knowing what you have and knowing how you got there, and it is the thing an inventory system is usually missing.
lending_service.py at 10 KB is the domain-specific part: carpets go out to contractors and come back, which is neither a sale nor a stock transfer and needs its own vocabulary.
An Excel export service, because the business this serves runs on spreadsheets and a system that cannot hand work back to a spreadsheet does not get adopted.
A GitHub Actions build workflow, the earliest CI in the personal repositories.
Trade-offs
- A service layer at this scale is more files than a small Flask app needs. It is what makes order and payment logic testable and reviewable in isolation, and it is the reason the largest module is a service rather than a view.
- Transaction-log stock over a quantity field: every read costs more, and stock history becomes auditable rather than reconstructed.
- Excel export is unglamorous integration work that is frequently the deciding factor in whether an internal tool is used at all.
2025-10 → 2026-03Model search, and the dashboard for reading it
A config-driven XGBoost hyperparameter pipeline that records every run with its SHAP analysis, and a React dashboard built from the same visual vocabulary.
- Repos
- MLops, private, 3 commits
- crypto-interface, private, 62 commits
- Stack
- Python · XGBoost · SHAP · Plotly · Parquet · React · Vite
Architecture
The pipeline is driven entirely by one config.yaml, paths, hyperparameter ranges, everything, and every run writes its configuration, metrics and output paths into a single JSON file acting as a local document store. The README is explicit that this is a deliberate stand-in for MongoDB, which is the correct instinct: pick the storage shape now, defer the storage engine.
Interpretation is part of the pipeline, not a follow-up. shap_analysis.py runs automatically for every trained model and persists the summary plot and feature-importance data next to the run. A search that records only scores tells you which configuration won; this one tells you what the winner is paying attention to.
Results are read through a Plotly parallel-coordinates plot, the correct chart for high-dimensional hyperparameter sweeps, where the question is which regions of the space are good rather than which single point is best.
Training data is standardised into Parquet by a separate preparation script, so the training loop never parses a CSV.
The React dashboard is built from the same primitives, a parallel-coordinates plot with brushing, SHAP plots, a dataset selector, a simulation chart with filters and metrics, and a chatbot panel, over a baked model-data JSON file rather than a live API.
Trade-offs
- Random configuration sampling, named "dumb grid" in the code, which is an honest name, rather than Bayesian optimisation. It parallelises trivially, it cannot get stuck, and the parallel-coordinates view is far more informative over a broad random sample than over a sequentially narrowing one.
- A JSON file as the experiment database: no concurrent writes, no query language, and zero infrastructure with a migration path already identified.
- SHAP on every run costs real time per model in exchange for interpretability that is present when you need it rather than regenerable in principle.
- A static baked JSON behind the dashboard means it deploys anywhere as a static site and goes stale the moment the pipeline runs again.
What is not claimed hereThese are two separate repositories that share a visual vocabulary, parallel-coordinates and SHAP plots over model runs. I have not verified that the dashboard reads this pipeline's output, so no such link is claimed. crypto-interface's own README is the unmodified Vite starter template and describes nothing.
2025-10LLM trading ledger
A model proposes trades against live price data; a separate ledger module is the only thing allowed to record them.
- Repos
- crypto-llm, private, 5 commits
- Stack
- Python · LLM API · GUI
Architecture
Five modules with one boundary that matters: data_fetcher gets prices, llm_handler reasons, trading_ledger records, persistence stores, price_updater runs on its own cadence. The model never writes to the ledger directly.
A ledger as a distinct module rather than a list of trades in the model loop makes position and history the system of record, and reduces the model to something that produces proposals.
Settings are isolated in config/settings.py and there is a create_structure.py scaffolding script, a habit visible across several of these repositories, where the project layout is itself generated.
Trade-offs
- A desktop GUI rather than a web interface: no deployment, no auth, no server, and it only runs where it is installed.
- A separate price updater decouples market data from the reasoning loop, at the cost of a second thing that has to be running.
- This is a small, early exploration of a pattern, model proposes, deterministic component decides and records, that shows up much more rigorously later in the retrieval and disaster-response work.
2025-07Offensive security attack chain
A four-stage chain from ARP poisoning to a reverse shell inside a repackaged Android app, written for a university security course.
- Repos
- Stack
- Python · Scapy · C · Android NDK · Smali · mitmproxy · BeEF
Architecture
Coursework for a cyber security module, built as a chain rather than a set of separate exercises: establish a machine-in-the-middle position by ARP spoofing both the victim and the gateway, use that position to answer DNS queries for a chosen domain with an attacker-controlled address, deliver a payload, and optionally inject a browser hook into unencrypted traffic with mitmproxy.
The payload stage is the substantial one, a native reverse shell compiled with the Android NDK, injected into a legitimate APK by decompiling it, adding the binary and Smali code to invoke it, then recompiling and signing. The repository contains both the C source and the Smali fragment.
The ARP spoofer restores both ARP tables in a finally block on exit. In a lab exercise that is a small courtesy; as a habit it is the same instinct as backup verification and key rotation, leaving the system in a state that does not require someone to come and fix it.
Trade-offs
- Two DNS spoofer implementations are kept in the repository, the second roughly four times the size of the first, the restart-rather-than-replace pattern again.
- Native C payload via the NDK rather than a framework-generated one: more work, and it demonstrates the actual mechanism rather than the tool that hides it.
- Written for a controlled lab environment against machines belonging to the exercise. It is included here because understanding the attack chain is what the hardening work on the other side of this page is built on.
2024-09 → 2024-12Autograders
Three repositories in one month, September 2024. The first code written to do a job while I was not present.
- Repos
- Stack
- Node.js · Express · React · Vite · Docker Compose · nginx
Architecture
A submission service and a front end, containerised together. An Express backend exposes a submit route; a React and Vite front end is served by nginx from its own image; Docker Compose brings up both. The whole thing is small on purpose, a grader that is hard to run does not get run.
Two near-identical repositories exist because the first was abandoned three commits in and the second carried on to thirty-four. The pattern that shows up later across five assistant repositories and three client-site repositories starts here.
Programming1 is the course side of the same work: a React site with a JSON-driven calendar and per-course content. Teaching material and the machine that marks it, built in the same month.
An exclude.txt in the backend, the grader is expected to ignore parts of a submission, which is the first sign of the problem being messier than "run the tests".
Trade-offs
- Containerising a student-facing tool early costs setup time and removes the class of problem where a grader works on the author's machine only.
- Two repositories rather than a rewrite in place: the abandoned one is still there, which is why the month reads accurately.
- A separate nginx image for a small front end is more infrastructure than needed, and it is the same shape as the production deployments that come two years later.
Before the account, pushed 2024-03Learning_python
Eighteen files recovered from a backup. The pattern that runs through everything else is already in them.
- Repos
- Stack
- Python
Architecture
A GPA calculator, two independent timetable generators, a binary encoder and a binary decoder, an xlsxwriter spreadsheet script, a Wi-Fi password reader, a text mangler. Then the games: hangman, tic-tac-toe, Ludo, rock-paper-scissors. Roughly 26 KB of Python.
Two of the files solve a problem that another file already solved. Longest_run_nSquared.py names its own complexity in the filename and sits next to longest run.py; TimeTableCreator.py sits next to timetable creator.py. Neither pair is a refactor, they are separate attempts, kept.
The repository description says the work predates the account and survived only because it was on OneDrive, and its README says plainly that the code is unoptimised and may not run.
Trade-offs
- Included here because it is evidence rather than a portfolio piece. The two habits that everything else on this page is built from, automate the chore, and rebuild rather than abandon, are both visible before there was any version control to record them.
2024-03 → 2024-05Games
Three engines in three months, the first work that produced something other than text.
- Repos
- Stack
- Python · Pygame · Godot · C# · Tiled
Architecture
A Doodle Jump implementation in Pygame from a course skeleton, with custom art and fonts, decoy and jumpy platform variants, two enemy types and a persisted high score.
Two Godot projects with C# scripting: a platformer with coins, a kill zone and a game manager, and a 2D portal game at thirty-five commits whose levels are authored in Tiled and imported as tilesets.
Three months, and the through-line is the toolchain rather than the games, going from a library you call, to an engine with a scene tree, to an engine plus an external level editor.
Trade-offs
- Course-provided skeletons rather than from-scratch engines: the work is in the gameplay and the integration, not the rendering loop.
- Choosing C# in Godot over GDScript aligned the game work with the C# used elsewhere at the time, at the cost of the engine's better-documented path.