Interview

From Models to Systems: Moxank Patel on Building Machine Learning That Works at Scale

IndustryTrends

Moxank Patel, Machine Learning Engineer at Meta, has worked across machine learning, information retrieval, NLP, computer vision, backend engineering and production systems. In this conversation, he discusses the engineering challenges that emerge when machine learning moves beyond experiments and into real-world applications.

Your career spans machine learning, software engineering, information retrieval and recommendation systems. How did that journey take shape?

My journey into machine learning was not a straight line. It started with a broader foundation in computer engineering and software development, and gradually moved toward solving problems where machine learning could make a measurable difference. 

I completed my bachelor's degree in Computer Engineering at Dharmsinh Desai University and later pursued a master's degree in Computer Engineering at San Jose State University. During that period, I became increasingly interested in problems involving large amounts of information and how software could process that information more intelligently. 

One early example was DocFinder, where I worked on search across more than 100,000 documents. We used a microservice architecture and locality-sensitive hashing to make retrieval more efficient. 

Looking back, that project taught me something that has stayed relevant throughout my career: building a model is only one part of the problem. You also have to think about how information reaches the model, how quickly the system responds and how the technology fits into the larger application. 

That thinking eventually led me toward machine learning and, ultimately, large-scale recommendation systems. 

What did working on information retrieval teach you about today's AI systems?

Retrieval is often the step people don't notice, but it is fundamental. 

If a system has millions of possible pieces of information, it cannot treat every item as equally relevant. It first needs to narrow the search space and identify useful candidates. 

That was the basic challenge behind DocFinder. We were working with a much smaller collection compared with today's large-scale AI systems, but the underlying problem was similar. 

Modern recommendation systems operate at a much greater scale. They may have millions of possible items to consider before ranking the most relevant candidates. So the question becomes not just "Can the model rank this?" but "How do we efficiently get the right information to the model in the first place?" 

That connection between retrieval and modeling is something I find particularly interesting. 

You have also worked with NLP and computer vision. What did those projects add to your understanding of machine learning?

They gave me exposure to very different types of data and problems. 

In YourDoc, I worked on processing information collected from medical websites and extracting symptoms from conversations. That involved preprocessing the data before applying natural-language-processing and neural-network techniques. 

I also worked on DeMystify, a computer-vision project focused on converting grayscale images into colour using a generative adversarial network. The model was trained on more than 100,000 images. 

These projects were technically different, but they taught me the same broader lesson: machine learning becomes valuable when it helps software deal with information that is difficult to capture through traditional rules alone. 

Documents, conversations, images and behavioural signals all require different approaches. Understanding those differences is important when designing a system that has to work outside a controlled research environment. 

What changes when you take a machine-learning model from an experiment and put it into a product?

The problem becomes much bigger. 

In a research environment, you are often focused on whether the model works. In production, you have to ask many more questions. 

How will the application communicate with the model? How will data reach it? How quickly does it need to respond? What happens when something fails? How do you monitor it? How do you update it? 

At Sequel, I worked on integrating OpenAI's language model into an existing application. I implemented more than five Python APIs and connected the functionality to a React/Next.js interface. 

The language model itself was not the entire project. The real engineering work was in connecting the model to the rest of the application so that users could actually make use of its capabilities. 

That experience reinforced the idea that AI is increasingly becoming another component of software rather than something that sits separately from it.

What are some of the practical bottlenecks you have encountered while building these systems?

There have been quite a few, and they are not always model-related. 

At Sequel, for example, I worked on data extraction from more than 100 online sources and contributed to increasing extraction speed by roughly three times. 

I have also worked on distributed services where improvements reduced server response time by around 40%, while changes to frequently accessed data contributed to another 30% performance improvement. One request-processing implementation was tested with 10,000 concurrent users. 

Another project involved reducing application memory consumption by 90%. 

These experiences showed me that performance problems can appear at almost any layer. Sometimes the bottleneck is the model. Sometimes it is the data pipeline, API, database, memory usage or infrastructure around it. 

Where does machine learning make the biggest difference in these situations?

I think the strongest use cases are those where machine learning addresses a genuine operational problem. 

For example, I worked on automating table-structure recognition in medical documents. The objective was to reduce the amount of manual effort involved in interpreting those documents. 

The resulting work reduced manual processing by 28%, while improvements to image processing reduced processing time by 53%. 

Those numbers matter because the purpose of the model was not simply to demonstrate that a neural network could recognize a table. The objective was to improve an existing workflow. 

That distinction is important. A technically impressive model does not automatically create a useful product. The real test is whether it improves the system around it. 

You have also worked on legacy applications and reliability. Why is that relevant to machine learning?

Because reliability becomes more important as systems become more interconnected. 

Earlier in my career, I worked on applications at the Institute for Plasma Research that involved older software and changing computing environments. Modernizing those applications and improving their stability contributed to a 41% reduction in downtime. 

At first glance, that may not look like machine learning work. But the same engineering principles apply to ML systems. 

A recommendation model can be highly accurate, but if the service is unavailable, if data cannot reach it or if the system takes too long to respond, the model's accuracy does not help the user. 

Reliability, latency and resource efficiency are therefore part of the overall AI engineering problem. 

How does your earlier engineering experience connect with your current work at Meta?

My current work at Meta involves large-scale recommendation models, and recommendation systems bring many of these problems together. 

You have large amounts of data, constantly changing information, retrieval, ranking, model inference and strict requirements around latency and infrastructure. 

My earlier work gave me experience with different parts of that broader system. I have worked on retrieval, NLP, computer vision, APIs, data processing, application performance and production deployment. 

The technologies change, but the underlying engineering questions remain surprisingly similar. 

How do you find the right information? How do you process it efficiently? How do you make the model useful to an application? And how do you make sure the entire system continues to work as demand increases? 

Do you think the distinction between software engineering and machine learning engineering is becoming less clear?

Yes, and I think that is a natural development. 

Machine learning used to be discussed primarily in terms of models and metrics. Today, a model is often one component inside a much larger software system. 

You have data pipelines, retrieval systems, APIs, databases, model-serving infrastructure, monitoring and user interfaces surrounding it. 

Recommendation systems make this particularly clear. Retrieval, ranking and serving cannot always be designed independently because changes in one layer can affect the others. 

As AI becomes more deeply integrated into products, engineers need to understand both the intelligence layer and the systems that make that intelligence usable.

What do you see as the biggest challenge for machine learning as systems continue to scale?

I think the challenge is moving from technically capable models to dependable systems. 

As models become more powerful, the surrounding infrastructure has to keep up. Data needs to move quickly. Systems need to be efficient. Inference needs to meet latency requirements. Applications need to remain reliable. 

The industry is already dealing with these issues at a very large scale. 

For me, the interesting part is that the question has remained consistent throughout my career, even though the scale has changed: 

What is preventing this system from delivering its intended value, and how can we remove that constraint? 

Sometimes the answer is slow data. Sometimes it is excessive resource consumption. Sometimes it is manual work or unreliable software. 

At scale, you often have to solve several of those problems at once. That is where machine-learning engineering becomes much more than model development. 

Crypto Prices Today: Bitcoin Slips to USD 84,200 as ETF Outflows and Liquidations Test Support

KuCoin vs Crypto.com: Fees, Features, Crypto Trading Options Compared

How Businesses are Integrating Cryptocurrency and Blockchain into Operations

Crypto Market Today: Regulation, Tokenized Funds, Binance Scrutiny

How Ethereum Staking and its Burn Mechanism Work