Lab — Testing the Edges of LLMs
AI Experiments
Every project here starts as a question about what LLMs can actually do, not what the marketing claims. I build production apps, agent systems, and prototypes to find the real edges — where a model reasons reliably, where it fails with confidence, and where the harness around it matters more than the model itself. Contact me to see these in action.
The operating system
An agentic research & operations vault
A connected research and operations vault wired through live data — meeting-transcript ingestion, external signal monitoring, and a structured knowledge base — that automates and augments PH1's research, synthesis, and client-deliverable workflows end to end.
Tools I've shipped & am building
AI Experiments
- Production · Live
AI value diagnostic app
A web app that helps enterprise leaders identify where their AI investments are failing at the behavioral layer — turning the PH1 methodology into a self-serve diagnostic.
- Production · Live
Publishing & content pipeline
The full infrastructure behind Product Impact — semantic search, automated publishing, social generation, and structured data — running as connected agents.
- Prototype · In development
AI harness built on reviewer concept
A harness built on a "reviewer" concept — giving teams a model-agnostic path to adopt LLMs while defining clear, precise review rules under ZDR (zero data retention), so the review logic stays portable across models even as the underlying provider changes.
- Prototype · In development
Site evaluation agent
A crawler-plus-analysis agent that assesses a site's readiness for the shift in how Google and LLMs route traffic to revenue.
- Prototype · In development
Travel planning agent
A planning agent that turns loose intent into a structured, bookable itinerary — a testbed for multi-step orchestration and tool use.
- Prototype · In development
Insurance selection tool & more
Plus context-graph integrations and an image-generation tool — a rotating set of experiments that each sharpen a capability I then reuse across the practice.
The methodologies behind the work
Frameworks
- PH1 · Incubate Practice
AI Product Calibration
A structured evaluation built for how AI actually fails — silently, not loudly. Every product is scored across four dimensions: Power (does it do what it claims), Speed (does it reduce effort or just move friction elsewhere), Impact (does it change behavior that matters), and Joy (does it earn enough trust to be recommended).
- PH1 · Every Engagement
The PH1 Method
The four-stage process behind every PH1 engagement: Baseline (pinpoint value drivers and experience gaps), Benchmark (build a truthful measurement framework against competitors), Build (produce the evidence for what to improve), and Transform (deliver compounding improvements).
- PH1 · Accelerate Practice
Digital Acceleration Pillars
The framework governing every Accelerate engagement — four shifts that separate organizations genuinely transforming from those that are just spending: Value (output to impact), Voice (broadcast to dialogue), Velocity (speed to ease), and Vision (isolated KPIs to unified purpose).
- PH1 · Service Design
Service Transformation Methodology
A four-stage approach to transforming complex services: Bridge Divides (map the stakeholder ecosystem so every voice is heard), Explore Possibilities (research what's possible, not just what's broken), Test Transformations (rapid prototyping and strategic foresight reveal what's viable), and Maximize Impact (service blueprints and training workshops make the change stick).
- PH1 · Conversion Research
Competitive Benchmarking
Business success is defined by how customers first interact with your product and its core features. This methodology quantifies the core conversion experience against your competitors — pinpointing where drop-offs are happening and where the clearest opportunities exist.
Why it matters
Building is how I test the edges
Every project here doubles as a controlled experiment — a way to find where a model's reasoning holds up, where it quietly breaks, and where the harness around it (rules, retrieval, review) matters more than the model itself. That judgment, earned by building rather than reading about it, is what I bring into client work.