AI Product Analyst

London, SE1 8NWPosted: 11th September 2026

About the role

We are building AI products for our clients and for colleagues across the group, in tax, HR, health and safety, and care. We can see that people use them. What we cannot yet see with any precision is how well they perform — for which clients, in which sectors, on which kinds of question — and therefore where to invest next.

This role answers that. You will own product analytics for our AI products: measuring performance across the dimensions that matter, building the reporting and dashboards the team and the business work from, and drawing out the insight that sits behind the figures. The purpose is to give the business a reliable evidence base for its decisions about where to invest.

You will also own the evaluation programme. Where analytics establishes how the product performs in the hands of real users, evaluation establishes whether a change is good enough to release. You will own the datasets the product is tested against and the analysis of what the results mean, working closely with the AI Domain & Workflow Specialist, whose professional judgement determines what a correct answer looks like.

The two areas are connected: findings from production inform the scope of evaluation coverage, and evaluation results inform how production performance should be interpreted.

What you will do

Product performance and analytics

•     Measure how our AI products perform across the dimensions that matter — retrieval quality, correctness of output, and engagement with what is produced.

•     Own production quality observability: thumbs-down rates, regeneration rates, task abandonment, and drift in how the product is being used.

•     Specify what needs to be measured and work with the engineers to get it captured, where the product itself has to change.

•     Build and maintain the analysis and dashboards the team works from day to day.

Insight and business decisions

The value of this role lies in the conclusions drawn from the data, not in the volume of reporting produced.

•     Interrogate the data to establish what is actually happening: identify patterns, outliers and probable causes, and determine which of them are material to the business.

•     Select the analytical approach that fits the question — segmentation, cohort comparison, trend analysis, correlation, or a purpose-built piece of work — rather than reporting a fixed set of measures.

•     Where segmentation is the appropriate method, cut the data by dimensions such as client, client type, subject matter, industry and question type, so that patterns are visible rather than lost in an average.

•     Translate findings into recommendations the business can act on: where the product is strong, where it is weak, where further investment is justified, and what should stop.

•     Present insight to the Product Manager to shape priorities and to the Director to inform the roadmap, and follow it through so that findings change decisions rather than accumulate in reports.

•     Establish whether the product is working, and for whom, on the evidence — including where the conclusion is unwelcome.

The evaluation programme

•     Own the evaluation programme for our AI products: what gets tested, how broad the coverage is, and what the results actually mean.

•     Analyse evaluation output — trends, regressions and gaps in coverage — and report it to the team in a form they can act on.

•     Work with the engineers who build and maintain the evaluation harness. They own the harness and its integration into the build pipeline; you own what it tests and what the output tells us.

•     Say clearly when coverage is not strong enough to support a release decision. You will not set the release gate — that sits with engineering and the Delivery & QA Manager — but your analysis is what informs it.

Golden datasets

The product is tested against curated sets of questions and known-good answers. Building and maintaining them well is one of the most valuable things this role does, and one of the most time-consuming.

•     Own those datasets: their construction, coverage, versioning and upkeep.

•     Run the elicitation with subject-matter experts — principally the AI Domain & Workflow Specialist, and advisers across the business — to capture the professional judgement behind a correct answer.

•     Design coverage deliberately: which question types, which sectors, and which difficult edge cases the product must handle well.

•     Keep the datasets current as legislation, guidance and professional practice change.

Reporting

•     Report internally to the AI team: the analysis the team acts on week to week, and the trends it needs to see early.

•     Provide the underlying analysis that the Delivery & QA Manager uses when reporting quality and delivery to the wider business, so that internal and external accounts tell the same story.

What we are looking for

Essential

•     Substantial experience as a product, data or commercial analyst, having owned measurement for a product or service end to end.

•     Strong hands-on SQL and Python, or close equivalents, and the ability to build your own analysis and dashboards without waiting on someone else.

•     The analytical judgement to select the right method for a question and to identify which lines of enquiry are worth pursuing, rather than producing measures because they are straightforward to produce.

•     The ability to turn analysis into a clear recommendation, and to hold that recommendation in front of senior people.

•     Comfort working with the output of AI or machine-learning systems, and an understanding of why measuring them differs from measuring a deterministic product.

•     Experience drawing knowledge out of subject-matter experts and turning it into something structured and reusable.

•     The willingness to report that something is not working, with the evidence to support it.

Desirable

•     Experience of AI or large language model evaluation — golden datasets, scoring, regression testing.

•     Familiarity with retrieval-augmented systems and how retrieval quality is assessed.

•     Experience in tax, HR, health and safety, or care, or in another regulated professional services environment.

•     Experience of experiment design and A/B testing.

•     Experience working alongside engineers where analysis feeds a build pipeline.

Why this role

•     You will define how a growing product function measures itself, rather than inheriting someone else’s metrics.

•     The evaluation half of the role means your work shapes what ships, not just what gets described afterwards.

•     You will work directly with in-house professional experts, on products where correctness genuinely matters to the person on the other end.

 

This job description sets out the main duties of the post at the date it was written. These may change as the role and the product develop, in consultation with the postholder.

Interested in this role?

Grouprecruitmentuk@peninsulagrouplimited.com

Find us on: