arpy dragffy

Lab — Testing the Edges of LLMs

AI Experiments

Every project here starts as a question about what LLMs can actually do, not what the marketing claims. I build production apps, agent systems, and prototypes to find the real edges — where a model reasons reliably, where it fails with confidence, and where the harness around it matters more than the model itself. Contact me to see these in action.

The operating system

An agentic research & operations vault

  • Transcript Ingestion
  • Signal Monitoring
  • Knowledge Base

A connected research and operations vault wired through live data — meeting-transcript ingestion, external signal monitoring, and a structured knowledge base — that automates and augments PH1's research, synthesis, and client-deliverable workflows end to end.

Tools I've shipped & am building

AI Experiments

  • AI value diagnostic app

    A web app that helps enterprise leaders identify where their AI investments are failing at the behavioral layer — turning the PH1 methodology into a self-serve diagnostic.

  • Publishing & content pipeline

    The full infrastructure behind Product Impact — semantic search, automated publishing, social generation, and structured data — running as connected agents.

  • AI harness built on reviewer concept

    A harness built on a "reviewer" concept — giving teams a model-agnostic path to adopt LLMs while defining clear, precise review rules under ZDR (zero data retention), so the review logic stays portable across models even as the underlying provider changes.

  • Site evaluation agent

    A crawler-plus-analysis agent that assesses a site's readiness for the shift in how Google and LLMs route traffic to revenue.

  • Travel planning agent

    A planning agent that turns loose intent into a structured, bookable itinerary — a testbed for multi-step orchestration and tool use.

  • Insurance selection tool & more

    Plus context-graph integrations and an image-generation tool — a rotating set of experiments that each sharpen a capability I then reuse across the practice.

The methodologies behind the work

Frameworks

  • AI Product Calibration

    A structured evaluation built for how AI actually fails — silently, not loudly. Every product is scored across four dimensions: Power (does it do what it claims), Speed (does it reduce effort or just move friction elsewhere), Impact (does it change behavior that matters), and Joy (does it earn enough trust to be recommended).

  • The PH1 Method

    The four-stage process behind every PH1 engagement: Baseline (pinpoint value drivers and experience gaps), Benchmark (build a truthful measurement framework against competitors), Build (produce the evidence for what to improve), and Transform (deliver compounding improvements).

  • Digital Acceleration Pillars

    The framework governing every Accelerate engagement — four shifts that separate organizations genuinely transforming from those that are just spending: Value (output to impact), Voice (broadcast to dialogue), Velocity (speed to ease), and Vision (isolated KPIs to unified purpose).

  • Service Transformation Methodology

    A four-stage approach to transforming complex services: Bridge Divides (map the stakeholder ecosystem so every voice is heard), Explore Possibilities (research what's possible, not just what's broken), Test Transformations (rapid prototyping and strategic foresight reveal what's viable), and Maximize Impact (service blueprints and training workshops make the change stick).

  • Competitive Benchmarking

    Business success is defined by how customers first interact with your product and its core features. This methodology quantifies the core conversion experience against your competitors — pinpointing where drop-offs are happening and where the clearest opportunities exist.

Why it matters

Building is how I test the edges

Every project here doubles as a controlled experiment — a way to find where a model's reasoning holds up, where it quietly breaks, and where the harness around it (rules, retrieval, review) matters more than the model itself. That judgment, earned by building rather than reading about it, is what I bring into client work.