I ship LLM systems into real customer-facing production — not demos. Sales agents, real-time voice AI, autonomous agents, and the self-hosted GPU infra underneath them.
📞 Call the number above and you'll talk to an AI receptionist I built — it screens the opportunity and texts me. The call itself is a demo of my work.
Claude sales closer on a live Shopify store + Facebook Messenger. Live cart awareness, session memory, dynamic product UI, and a hallucination guard that blocks fabricated prices before they reach a customer. Real paying customers.
Real-time voice agent: Deepgram STT → streaming LLM → ElevenLabs TTS, low-latency per-sentence. Caller memory, 6 industry personas, ROI/booking tool calls, missed-call text-back.
llama.cpp serving a 27B model across a Tesla M40 + RTX 3060. Custom CUDA build for Maxwell hardware; tuned to 19.97 tok/s (1.90× over naive split) by benchmarking layer placement.
An autonomous agent that health-checks systems, ranks marketing opportunities, and publishes content on a schedule with no human in the loop. First autonomous post verified in production.
One-shot AI business launcher: intake form → website + 3-variant ads + 7-day social calendar + video commercial + jingle, for ~$1.40 per client.
Food-delivery platform (Laravel + Flutter, 183 restaurants, AI menu photos, voice ordering) and a live client marketing site with AI imagery + Kling hero video.
Marion, OH. Built and operate the entire technical stack: live storefront, AI sales agent, voice AI, autonomous marketing, self-hosted PBX, payments, and self-hosted multi-GPU LLM inference. Every project above is deployed here. Real customers, real revenue.
I go into the messy real problem, wire the model into the actual workflow — CRM, storefront, phone system, payments — and put guardrails on it so it's safe in front of a paying customer. That's the Forward Deployed / AI Automation Engineer job.