Machine Learning Manager
Jakub leads the AI engineering team at LexisNexis Intellectual Property Solutions. This is the team behind Protégé's agentic patent analysis system. His background spans data science and engineering roles at companies including Google/YouTube, Essentia Analytics, and Trilodocs, and he holds an MSc in Data Science and AI from the University of London.
Welcome to the fourth article in our series, “Conversations With the Team That Built Protégé in PatentSight+.” In the first post, we discussed how we manage hallucination risk by constraining the agentic system. In the second, we explored how transparency lets users inspect the analysis as it develops. And in the third, we looked at why patent analysis benefits from a hybrid of structured workflows and autonomous planning, rather than a fully deterministic or fully unconstrained system.
This post turns to the people behind the specific decisions mentioned in the above articles: the AI engineering team that builds and maintains LexisNexis Protégé™ in PatentSight+™.
Developing and improving Protégé in PatentSight+ is a team effort, and the team’s boundaries evolved to include professionals across multiple domains:
In this blog post, we write about the team of AI engineers who worked specifically on Protégé in PatentSight+. Looking at things like day-to-day functions, how they validate their work, and how they collaborate with others.
Software development is a broad domain with many specializations. Some software engineers specialize in building the user interfaces seen in applications and on websites (often called frontend engineers); others focus on processing and storing data (data engineers); and others design the systems between these boundaries (backend engineers). Compared to them, AI engineers operate in a much younger domain.
The title “AI engineer” became more widely used after 2022, following the launch of OpenAI’s ChatGPT. It proliferated as an umbrella term for several responsibilities. AI engineers are typically backend engineers specializing in AI systems. But, depending on the organization, they may also need to contribute to user interfaces and data platforms.
In our team, an AI engineer’s role is that of a software developer crossed with a data scientist. As engineers, they need to be able to build, maintain, and extend both backend and frontend systems; as data scientists, they need to understand frontier model architectures and rigorous statistical testing well enough to confidently evaluate an AI product’s performance. Admittedly, this is a somewhat rare combination of skills to find, which brings us to our next topic.
Even “AI” itself refers to different things across the industry, so the hiring process always starts by agreeing on what it means to be an engineer responsible for an AI system.
We front-load the hiring process with an alignment exercise that includes a small, live-coding data science test (e.g., transposing a matrix or implementing a sigmoid activation function). It is not enough to ask a candidate whether they possess a prerequisite skill (to an applicant, all such questions sound like, “Do you want this job?”). Nor is it sufficient, in the age of AI, to verbally quiz someone. Even if the majority of the code they produce during their engagement will be written with AI assistance, we still want to see them write code by hand during the interview because we believe that good judgment and expertise come from real internalized experience.
Later on, there is a larger coding exercise designed to test whether the engineer has the right intuitions about AI systems, can work within and extend an existing codebase, and can articulate their reasoning for choosing one data science algorithm over another. We have seen that even engineers with many years of experience often fail at this stage – the required combination of engineering prowess and data science skills is rare and elusive, and competition for the talent is at an all-time high. We acknowledge this scarcity and attempt to address it by offering engineers learning opportunities through a curated training curriculum designed to close expertise gaps, provided they demonstrate a sufficient baseline of skills throughout the hiring process.
While developing Protégé in PatentSight+, we have established standard processes to address the needs of the team, the product, and the customers. These processes are owned and led by the AI engineers directly, and they include (but are not limited to):
Taken together, these processes are designed around the core idea that an AI product’s success is driven mostly by the people and teams surrounding it. AI engineering remains a nascent and often under-defined discipline. Defining it well, hiring the right team members, designing operational processes, and creating measurable standards are what allow a team to ship a useful AI product that customers can trust.
Across this series, we’ve tried to show the thought process that actually goes into developing an agentic AI system that IP professionals can stake real decisions on: constraining the agent to manage hallucination risk, designing for transparency so users can see the reasoning behind every answer, and choosing an architecture that balances structure with autonomy instead of forcing a choice between the two. This final post is the piece that ties the rest together. All of the decisions come from a team that was deliberately hired and guided by processes built for exactly this kind of work.
If there’s one takeaway from the entire series, it is that an AI tool’s reliability depends less on the model beneath it and more on the people and processes put in place.
Schedule a demo to see how Protégé in PatentSight+ turns that engineering discipline into easily retrievable patent insights your team can confidently act on.
Customer Story
“Once people see the benefits, nobody will want to go back.” That’s how the analyst described using Protégé to cut days of portfolio analysis down to minutes. Learn what changed for them.