Articles

AI engineering notes from real workflow work

Practical writing on agentic systems, orchestration, automation design, and the operating patterns behind reliable AI delivery.

LL
LLM Infrastructure • Mar 20, 2026

Tooling Map: self-hosted inference

A map of useful tools, libraries, and platform decisions. how to serve open models with throughput, batching, and cost control

Read article → Save
LL
LLM Infrastructure • Mar 18, 2026

Team Workflow: self-hosted inference

How engineering, product, and operations teams should collaborate. how to serve open models with throughput, batching, and cost control

Read article → Save
LL
LLM Infrastructure • Mar 17, 2026

Deployment Playbook: self-hosted inference

How to prepare the workflow for CI/CD and production operations. how to serve open models with throughput, batching, and cost control

Read article → Save
LL
LLM Infrastructure • Mar 15, 2026

Security Review: self-hosted inference

The security and permission concerns that should be reviewed. how to serve open models with throughput, batching, and cost control

Read article → Save
LL
LLM Infrastructure • Mar 14, 2026

Evaluation Strategy: self-hosted inference

How to measure quality, reliability, and operational readiness. how to serve open models with throughput, batching, and cost control

Read article → Save
LL
LLM Infrastructure • Mar 12, 2026

Common Failure Modes: self-hosted inference

The mistakes teams should identify before launch. how to serve open models with throughput, batching, and cost control

Read article → Save
LL
LLM Infrastructure • Mar 12, 2026

LLM Gateways and Model Routing Are Becoming Core AI Infrastructure

Why teams need model routing, budgets, fallback logic, observability, and provider abstraction before AI usage scales.

Read article → Save
LL
LLM Infrastructure • Mar 11, 2026

Implementation Checklist: self-hosted inference

A practical checklist for building and reviewing the workflow. how to serve open models with throughput, batching, and cost control

Read article → Save
LL
LLM Infrastructure • Mar 09, 2026

Architecture Guide: self-hosted inference

A system-design view for planning production implementation. how to serve open models with throughput, batching, and cost control

Read article → Save