Since large language models started integrating directly into code editors, the shape of software consulting has changed in a way that benefits clients more than it threatens them. My daily loop is now human architecture + machine throughput: I design systems, own the judgment calls, and use LLMs to compress the hours between decisions.
A lot of “AI development” is marketing for cloud token bills. I run local-first AI infrastructure instead — llama.cpp serving GGUF models on consumer GPUs behind an OpenAI-compatible endpoint, wired straight into VS Code. Client data never leaves my machine during development, iteration is near-instant, and there's no per-token meter running while we build.
The same stack powers client-facing products: RAG chatbots over your own documents with hybrid retrieval (vectors + keyword + query rewriting), NL-to-SQL analytics, knowledge-graph traversal, and MCP endpoints so existing tools can talk to your data. When AI belongs in a product, I build it like plumbing — boring, testable, observable — not like a demo.
$ session --mode local
▸ architect: human (me)
▸ throughput: LLM-assisted
▸ models: GGUF @ localhost:8080
▸ data_egress: none during dev
▸ rag: pgvector + BM25 + rewrite
▸ nl_to_sql: yes
▸ mcp_endpoints: chat/analytics/health
# vram: 4GB. it works.