MathiasBoulanger
Product Manager, AI Product BuilderParis, relocating to Buenos Aires. Available remote.
I'm drawn to products that change how people work. On some I lead the team that builds them; on others I build them myself, directing AI.
The work changed shape
A product manager used to hand a spec to a team and wait. That wait is where most ideas die: by the time the thing exists, the question it answered has moved.
Now I scope the problem, direct the AI that builds the solution, review what comes out, and ship it. The loop that took a quarter takes an afternoon, and the hypothesis gets tested while it still matters.
It does not replace the team. It changes which questions need one.
Inside a company of a hundred people
Product ownership on one side, adoption on the other. The numbers below are the ones that were measured and reported, not the ones that sounded best.
- 20%
- of company revenue, the platform I own
- 100+
- people in the AI adoption program I run
- 3h→ 5min
- a manual triage, after one automation
- 94%
- customer verbatim coverage, turned into actionable items
- 5–10h/wk
- returned to a business development team
- 9markets
- and 7 languages the store ships to
Eleven things, three ways of working
Some of these I led, some I built with my own hands, and some I taught someone else to build. Filter by which.
Open a row for the evidence
Three programmes at once: a Symfony 1 to Symfony 7 migration, a headless CMS internal teams operate themselves, and the rollout of a new design system. Two squads across Design, Tech, CRM and Marketing.
Two company-wide training sessions, a champion per department, executive sponsors, and results presented to the executive committee. One attendee became the first measured case.
Two sessions delivered to the whole company. One attendee went on to become the first measured case: a mailbox triage that went from three hours to five minutes.
Trained a non-technical CEO to direct AI himself: what to delegate to a model, what to review, and the security practices around it. He shipped a full backend migration to Django in a week, work that used to sit in his development team's queue.
Ask a question in plain language, get an answer built on real product indicators and live satisfaction feedback, with sources cited. Runs on a local model so nothing leaves the company.

One persistent context that Product, Engineering and QA read from, instead of each briefing their own tools from scratch. Built on the Model Context Protocol so any client can consume it.
Every inbound lead classified across three profiles, returned with a summary and the next action to take. Presented as a live session published by No-code France.

Output went from 2 to 11 pieces a month: automated posting, article drafting and newsletter workflows. Published as a case session by No-code France.

An F1 prediction game with real players. Cron jobs pull results from the public OpenF1 API and settle scores with no manual input.

A conversational search agent over a podcast archive. Answers cite their episode and link to the exact second, and a labelled retrieval evaluation sets the relevance threshold.

Reads a post's metrics against the account's own median and explains what to change. Every claim names the metric behind it.

The part that separates a demo from a product
Anyone can get a model to answer once. The work is knowing whether it answered well, and noticing when it stops.
Evaluation on creative output
Text generated at temperature 1 defeats exact-string tests, so the golden set asserts on properties instead: valid shape, claims grounded in the source, a banned vocabulary, hypotheses phrased as hypotheses.
eval-prompts.ts, from the Reel Coach project
A retrieval threshold placed from data
Two labelled question sets, on-topic and off-topic. The gap between them is what tells you where the relevance floor goes. It caught a timestamp bug that had looked like a ranking problem.
audit-retrieval.ts, from the PodSearch project
Guardrails that refuse before the mistake
Fifty agent skills and a set of hooks that block a bad commit, a stray log or an edit to a secrets file, rather than catching it in review.
A personal OS, wired into mail, calendar, Slack, Jira and the knowledge base
Continuous integration, not a demo branch
Seven repositories run their tests on every push. A prototype that cannot survive its own pipeline is a screenshot.
GitHub Actions, across 7 repositories
What I'm looking for
A company building AI products, or one that needs someone to make AI land inside it. Remote, across European and American time zones.
