I don't think you can credibly write AI policy for an enterprise if you've never shipped anything with the tools yourself. These are the projects that keep me honest — real systems, with real users, that either work or don't. Each one taught me something I've since carried back into how we govern and deploy AI at work.
I wanted to explore one question: could conversational AI replace most of the traditional user interface of a calorie-tracking application? Every food-tracking app makes you search a database, pick a food, estimate a serving, navigate screens and record a meal. I wanted to take a picture and let the AI handle the rest.
The design requirement was frictionless mobile use. I built a custom Calorie Tracker GPT and pinned it on my phone. A typical interaction is:
From there the same interface handles everything else conversationally:
The spreadsheet is deliberately not the interface. It's the data layer. ChatGPT is how I interact with that data.
The system separates the conversational experience, the integration logic, and the persistent store — and data moves in both directions, so the same interface can both perform transactions and reason over history.
The GPT accepts a photograph or a natural-language description. For photos it identifies likely foods, estimates portions, estimates calories/protein/carbs/fat, distinguishes estimates from known values, and converts the result into structured data for the backend.
When real nutrition information exists — a package, a label, a restaurant, or something I provide — those values take priority over AI estimates. The goal was never to pretend visual estimation is precise. It was to make reasonable estimates and handle the uncertainty honestly.
| Metric | Daily target |
|---|---|
| Calories | 1,900 |
| Protein | 150 g |
| Carbohydrates | ≤ 150 g |
| Fat | ~70 g |
Calories and protein are the primary objectives; carbs and fat provide guidance. That lets the GPT give contextual advice instead of just displaying numbers:
After living with Version 1, I expanded the system past logging and retrieval into record-level actions, finer-grained data, historical analysis and visualization.
Every record now gets an automatically generated unique Entry ID from n8n. This turned out to be the most important architectural change in Version 2.
Without a unique identifier, an AI assistant can name something conversationally — "the protein shake I had this afternoon" — but the backend has no unambiguous way to know which row that is. Entry IDs are the bridge between natural-language intent and deterministic backend operations.
Version 1 stored a whole meal as one record. Version 2 stores identifiable foods individually — chicken, cauliflower rice and a protein shake become three records, each with its own Entry ID. The user never manages those records; the GPT still presents them conversationally as one meal.
The extra granularity makes individual items deletable and makes real analysis possible later. Composite dishes that can't reasonably be separated — chili, soup, casseroles — are still stored as one record.
Version 2 added a deleteFoodEntry workflow and action. Deletion is deliberately multi-step. When I say "delete the protein shake I just had," the GPT does not guess a row:
If more than one record could match, it asks for clarification rather than taking an ambiguous destructive action. Technically a small feature. Practically, the clearest lesson in the project about safe agentic design.
Once the model could perform transactions, the instruction set had to cover authority, not just tone. The system is instructed to:
That list is the real distinction between chatbot prompting and agentic system design. Once AI can act, the instructions have to define what it is authorized to do, under what conditions, and how it behaves when it isn't sure.
The retrieval API works by individual date, so for multi-day questions the GPT retrieves each date and combines the results. That supports "analyze my nutrition for the last seven days" — daily and average calories and protein, carbs and fat, days meeting targets, highest and lowest days, and general trend.
One data-quality rule matters more than the rest: a date with no records is missing data, not zero calories. Otherwise incomplete logging quietly improves your averages, and the system starts lying to you.
After retrieving the real records, the GPT runs its own data analysis and generates visualizations — daily calories against the 1,900 target, daily protein against the 150g target. That moves the application from a transactional logger to a lightweight analytics interface, with one natural-language surface responsible for the whole chain:
I stopped development after Version 2 on purpose — to actually use the system and find out which capabilities earn their complexity before adding more. That restraint is itself the lesson: the failure mode in personal AI projects isn't building too little, it's building features nobody, including you, turns out to need.
Microsoft's AI-900 (Azure AI Fundamentals) is a certification worth having on a team that's deploying Azure AI in production. Good practice exams are either expensive or bad. So I built one — not a prompt, not a chatbot, but an actual multi-user web application with accounts, a question bank, and a database behind it.
A web application where a user creates an account, logs in, and takes practice tests aligned to the Microsoft AI-900 exam objectives. Every answer is written to a backend database against that user's account, so attempts and progress persist across sessions rather than evaporating when the tab closes.
The calorie tracker demonstrates orchestration. This one demonstrates that I can stand up the whole shape of a product: authentication, a persistent multi-user data model, state that survives a session, a front end someone who isn't me can use, and a deployment that runs without my laptop open.
It was built using Google Antigravity — an agentic development environment — which made it a useful experiment in its own right. Building a full-stack application primarily by directing an agent is a genuinely different discipline from writing the code yourself, and it surfaced exactly the questions I now ask my own development teams about AI-assisted coding: where does review happen, what does the agent get wrong quietly, and what still requires a human to hold the architecture in their head.
A working lead-generation site for a luxury real estate practice, designed and built with Claude and ChatGPT — the positioning, the conversion copy, the layout, and the automation that runs underneath it. Unlike the other projects here, this one exists to produce business results for someone else, and it does.
Because the interesting part isn't the website. It's that one person, using AI tools deliberately, produced a complete marketing system — positioning, copy, design, integrations and automation — that would traditionally require a small agency and a meaningful budget. That's the productivity shift enterprises are trying to understand, demonstrated at small scale where I could see every part of it.
Visit the site →A blog covering AI tooling for technology leaders, running on an agentic content pipeline I designed, deployed and operate myself. It functions as a publication and, more usefully, as a live testbed for the agent patterns I'm evaluating for enterprise use.
Running an automated pipeline continuously is a different education from building a demo. Things degrade, sources go stale, output quality drifts, and you learn quickly which parts of an agentic system need a human gate and which genuinely don't.
Visit the site →Happy to go deeper on the architecture, the failure modes, or what any of it implies for doing this at enterprise scale.