MathiasBoulanger

Product Manager, AI Product BuilderParis, relocating to Buenos Aires. Available remote.

I'm drawn to products that change how people work. On some I lead the team that builds them; on others I build them myself, directing AI.

The work changed shape

A product manager used to hand a spec to a team and wait. That wait is where most ideas die: by the time the thing exists, the question it answered has moved.

Now I scope the problem, direct the AI that builds the solution, review what comes out, and ship it. The loop that took a quarter takes an afternoon, and the hypothesis gets tested while it still matters.

It does not replace the team. It changes which questions need one.

Inside a company of a hundred people

Product ownership on one side, adoption on the other. The numbers below are the ones that were measured and reported, not the ones that sounded best.

20%
of company revenue, the platform I own
100+
people in the AI adoption program I run
3h→ 5min
a manual triage, after one automation
94%
customer verbatim coverage, turned into actionable items
5–10h/wk
returned to a business development team
9markets
and 7 languages the store ships to

Eleven things, three ways of working

Some of these I led, some I built with my own hands, and some I taught someone else to build. Filter by which.

Open a row for the evidence

Three programmes at once: a Symfony 1 to Symfony 7 migration, a headless CMS internal teams operate themselves, and the rollout of a new design system. Two squads across Design, Tech, CRM and Marketing.

DirectusSymfonyJiraFigma

Two company-wide training sessions, a champion per department, executive sponsors, and results presented to the executive committee. One attendee became the first measured case.

EnablementTrainingGovernance

Two sessions delivered to the whole company. One attendee went on to become the first measured case: a mailbox triage that went from three hours to five minutes.

TrainingAdoption

Trained a non-technical CEO to direct AI himself: what to delegate to a model, what to review, and the security practices around it. He shipped a full backend migration to Django in a week, work that used to sit in his development team's queue.

EnablementDjangoSecurity practices

Ask a question in plain language, get an answer built on real product indicators and live satisfaction feedback, with sources cited. Runs on a local model so nothing leaves the company.

Local LLMRAGSQLFastAPI
The agent's question screen

One persistent context that Product, Engineering and QA read from, instead of each briefing their own tools from scratch. Built on the Model Context Protocol so any client can consume it.

MCPTypeScript

Every inbound lead classified across three profiles, returned with a summary and the next action to take. Presented as a live session published by No-code France.

n8nAirtableTallyLLM APIs
Watch the session
The scoring workflow in n8n during the live session

Output went from 2 to 11 pieces a month: automated posting, article drafting and newsletter workflows. Published as a case session by No-code France.

Maken8nClaude
Watch the session
The No-code France session on the Make automation

An F1 prediction game with real players. Cron jobs pull results from the public OpenF1 API and settle scores with no manual input.

Next.jsPrismaPostgresRailway
Play it
The Podium Fantasy race calendar on mobile

A conversational search agent over a podcast archive. Answers cite their episode and link to the exact second, and a labelled retrieval evaluation sets the relevance threshold.

pgvectorVoyageClaudeNext.js
Ask it something
The PodSearch entry screen

Reads a post's metrics against the account's own median and explains what to change. Every claim names the metric behind it.

Instagram APIClaudePrisma
Reel Coach analysing a post

The part that separates a demo from a product

Anyone can get a model to answer once. The work is knowing whether it answered well, and noticing when it stops.

  1. Evaluation on creative output

    Text generated at temperature 1 defeats exact-string tests, so the golden set asserts on properties instead: valid shape, claims grounded in the source, a banned vocabulary, hypotheses phrased as hypotheses.

    eval-prompts.ts, from the Reel Coach project

  2. A retrieval threshold placed from data

    Two labelled question sets, on-topic and off-topic. The gap between them is what tells you where the relevance floor goes. It caught a timestamp bug that had looked like a ranking problem.

    audit-retrieval.ts, from the PodSearch project

  3. Guardrails that refuse before the mistake

    Fifty agent skills and a set of hooks that block a bad commit, a stray log or an edit to a secrets file, rather than catching it in review.

    A personal OS, wired into mail, calendar, Slack, Jira and the knowledge base

  4. Continuous integration, not a demo branch

    Seven repositories run their tests on every push. A prototype that cannot survive its own pipeline is a screenshot.

    GitHub Actions, across 7 repositories

What I'm looking for

A company building AI products, or one that needs someone to make AI land inside it. Remote, across European and American time zones.

Mathias Boulanger