ABL-000 · SELF-HOSTED MIND · DESCENT VECTOR LOCKED

I build the AI systems you'd otherwise hire a team for.

Custom agents, workflow automation, and self-hosted LLM deployment — shipped as working software, not slideware. From the model weights up to the browser automation that drives the whole thing.

LOG // 1 CRAFT INBOUND · PILOT: A. HOSSAIN · DESTINATION: LAB

runs local multi-agent 128 tests / deploy 27B on 8 GB

29.9M

parameters pretrained from random initialisation — 2.0B tokens, 9.0 hours, one RTX 3070 Ti. Best val loss 1.341. Own architecture, own tokenizer, own weights.

27B

served on that same 8 GB card — CUDA llama.cpp from source, GGUF quantization, MTP self-speculative decoding. Measured, not guessed →

128

tests gate every deploy of the platform. 7,600+ records in the store, 450+ commits, one operator running it under pm2.

36/0/0+31

agents built / partial / scaffold, graded from source by a published rule → — 300+ lines and no stub markers, not self-assessment. The rule is on the page so the count can be checked. The +31 are specified and not built — each with the files it would touch and how it fails safe — counted separately, because a specification is not an agent.

§02

Selected Work

full dossiers →
01

REC 01 // MISSION RECORD

ABL — self-hosted multi-agent platform

A production personal-AI system coordinating many specialized agents behind one API. FastAPI backend routing across a fleet of locally run models, a vector-searchable store I schema-designed, and an egress boundary that classifies every outbound request and keeps sensitive data on the machine.

Python · FastAPI · local LLMs · RAG · pm2  ·  7,600+ records · 128 tests · 450+ commits (private repo)  ·  public excerpt: github.com/insomniac-asif/abl-core
02

REC 02 // MISSION RECORD

Modular trading & analysis system

Separated analytics, data, decision, execution, simulation, dashboard, and testing layers. Sole-authored the React 19 + TypeScript operator dashboard, plus SQL-backed ingestion and evaluation pipelines with logging and reporting.

Python · React 19 · TypeScript · Vite · SQL  ·  ~9,300 LOC · 43 modules  ·  full dossier →
03

REC 03 // MISSION RECORD

Self-hosted LLM serving & quantization

Running 27B–35B models on a single 8 GB consumer GPU: CUDA llama.cpp built from source, GGUF quantization, MTP self-speculative decoding, and MoE expert-offload tuning. Measured, not guessed.

llama.cpp · CUDA · GGUF · WSL  ·  read the write-up →
04

REC 04 // MISSION RECORD

Production Node.js services, live users

Discord bots on live servers: 80+ registered operations across 20+ modules, persistent per-user state, automated moderation, and an authorization layer with permission tiers that gates destructive operations behind explicit approval.

Node.js · authz tiers · state machines
§03

Stack

  1. 630.0 nm EDGE Cloudflare · Workers · Pages static · edge-cached
  2. 557.7 nm INTERFACE TypeScript · React · Playwright hand-built · no framework
  3. 486.1 nm SERVICE Python · FastAPI · Node.js reading telemetry…
  4. 427.8 nm DATA SQL / SQLite · Chroma · RAG reading telemetry…
  5. 391.4 nm RUNTIME llama.cpp · Ollama · ABLE reading telemetry…
  6. 777.4 nm METAL CUDA · Linux / WSL · Docker · RTX 3070 Ti reading telemetry…

Git throughout. Readings on the right are live where the lab can actually see them.

§04

Writing

TX 04.2
2026-08-20

I trained a language model from scratch on my gaming PC

29.9M parameters, my own architecture code and my own tokenizer, random weights to coherent prose in sixteen minutes — nine hours on one RTX 3070 Ti, including the two silent bugs that made the first attempt eighty-six times too slow.

TX 04.1
2026-08-20

What actually makes a 27B model faster on an 8 GB GPU

Six experiments on a single RTX 3070 Ti. Multi-token prediction doubled throughput for free, speculative decoding made things worse, and the biggest win was a Windows setting silently moving VRAM into system RAM.

all transmissions →

FIG.02 — SENSOR ARRAY · LOW ORBIT GROUND APPROACHING
PRESENCE = ABSENCE03/64

TOUCHDOWN · ALT 0 M  —  CRAFT SECURED · PILOT HOME

■ PRESSURE EQUALISED · ENTERING

The door was already open.

Nobody let me in. Nobody was here. The lanes had been running the whole time I was away.

ROOM 01 THE FLOOR one model, three personas, lit by whatever they are actually doing →
§05

Contact

The lab is running. Leave a transmission.

Tell me the problem in a sentence or two. If it's a fit I'll come back with a scope — usually same day.

Not here to hire? Everything else — TikTok, Discord, Hugging Face, and who I actually am — is on about & signals.

THIS PAGE · MEASURED IN YOUR BROWSER

REQUESTS
—
DOM NODES
—
FIRST PAINT
—
TRANSFERRED
—

Hand-written HTML, CSS and JavaScript — no framework, no build step, no component library. These are read from your browser when the page finishes loading, so they are yours, not mine: reload and they move with your connection and your cache.