ABL-000 · SELF-HOSTED MIND · DESCENT VECTOR LOCKED

I build the AI systems you'd otherwise hire a team for.

Custom agents, workflow automation, and self-hosted LLM deployment — shipped as working software, not slideware. From the model weights up to the browser automation that drives the whole thing.

LOG // 1 CRAFT INBOUND · PILOT: A. HOSSAIN · DESTINATION: LAB

runs local multi-agent 128 tests / deploy 27B on 8 GB
§02

Selected Work

full dossiers →
01

REC 01 // MISSION RECORD

ABL — self-hosted multi-agent platform

A production personal-AI system coordinating many specialized agents behind one API. FastAPI backend routing across a fleet of locally run models, a vector-searchable store I schema-designed, and an egress boundary that classifies every outbound request and keeps sensitive data on the machine.

Python · FastAPI · local LLMs · RAG · pm2  ·  7,600+ records · 128 tests · 450+ commits  ·  github.com/insomniac-asif/abl-core
02

REC 02 // MISSION RECORD

Modular trading & analysis system

Separated analytics, data, decision, execution, simulation, dashboard, and testing layers. Sole-authored the React 19 + TypeScript operator dashboard, plus SQL-backed ingestion and evaluation pipelines with logging and reporting.

Python · React 19 · TypeScript · Vite · SQL  ·  ~9,300 LOC · 43 modules
03

REC 03 // MISSION RECORD

Self-hosted LLM serving & quantization

Running 27B–35B models on a single 8 GB consumer GPU: CUDA llama.cpp built from source, GGUF quantization, MTP self-speculative decoding, and MoE expert-offload tuning. Measured, not guessed.

llama.cpp · CUDA · GGUF · WSL  ·  read the write-up →
04

REC 04 // MISSION RECORD

Production Node.js services, live users

Discord bots on live servers: 80+ registered operations across 20+ modules, persistent per-user state, automated moderation, and an authorization layer with permission tiers that gates destructive operations behind explicit approval.

Node.js · authz tiers · state machines
§03

Stack

Python · FastAPI · TypeScript · React · Node.js · SQL / SQLite · llama.cpp · Ollama · CUDA · RAG · Playwright · Linux / WSL · Docker · Cloudflare · Git

§04

Writing

TX 04.1
2026-08-20

What actually makes a 27B model faster on an 8 GB GPU

Six experiments on a single RTX 3070 Ti. Multi-token prediction doubled throughput for free, speculative decoding made things worse, and the biggest win was a Windows setting silently moving VRAM into system RAM.

all transmissions →

FIG.02 — SENSOR ARRAY · LOW ORBIT GROUND APPROACHING
PRESENCE = ABSENCE03/64

TOUCHDOWN · ALT 0 M  —  CRAFT SECURED · PILOT HOME

§05

Contact

The lab is running. Leave a transmission.

Tell me the problem in a sentence or two. If it's a fit I'll come back with a scope — usually same day.