Tag Archives: Dispatch

Powering AI-led research through simulation

Post Syndicated from Grab Tech original https://engineering.grab.com/powering-ai-led-research-through-simulation

The short story

Consider a Friday evening. A food order arrives from a mall in the city center. One driver is nearby; another is finishing a drop-off and will be available shortly; a second order from the same mall may or may not appear in the next two minutes. Dispatch the nearby driver now, or hold briefly for a batching opportunity? The decision window is short.

A fulfillment marketplace makes these decisions continuously. Each one is small. Across a city, those decisions determine whether your dinner arrives hot and whether a driver’s hour is well spent.

And that is one decision. There are dozens more: how far to look for a driver, when two orders are worth combining, which of three waiting trips gets the one free bike, how long to keep trying before giving up. These decisions interact, and the setting that is right for Friday at seven is wrong for Tuesday at two. Together they define a space of possible strategies far larger than anyone could exhaustively explore.

We sample only about a dozen points in that space each year. Not for want of ideas: implementing each strategy costs weeks of engineering, and a production experiment takes weeks more to judge. Some questions have no production answer at all. Nobody can run last Saturday again with twenty percent fewer drivers. We were never short of compute resources, and never short of data. We were short of attempts.

So we built sim-rs, a simulator that makes each attempt take minutes, and connected agents that run experiments on it.

sim-rs is self-contained: it requires no production services or databases. The dispatch lifecycle runs in one compiled program alongside a separate dispatch service that is spawned locally and operates offline. Historical booking and driver data go in as plain files, together with a configuration describing the strategy to test. Simulated bookings, trips, drivers, and a single metrics report come out. Processing one city-day of marketplace activity takes tens of minutes, from raw input to finished report. One command in, one report out: a workflow as practical for a software agent as for an engineer.

That compact contract is what makes the simulator AI-friendly. Agents can change bounded components, run reproducible experiments, receive verdicts they cannot alter, and help keep both production logic and behavioral models current.

Building blocks

The essential components

To model a marketplace, the simulator needs four core concepts:

  • Bookings: requests to move something (a passenger, a meal, a parcel), represented in one format regardless of the business vertical.
  • Drivers: simulated workers with a location, a shift, a vehicle, and a queue of work.
  • Trips: units assigned to drivers, containing one booking or several batched into a multi-stop route.
  • Ticks: simulated time, advancing in fixed steps. Nobody waits in real time.

At each tick, the simulator runs the following sequence:

  1. Demand arrives. Historical (or synthesized) bookings whose time has come enter the pool.
  2. Supply moves. Drivers come on shift, advance along routes, go idle, and reposition.
  3. Pre-dispatch cancellations are applied. Cancellation models, including survival-analysis models trained on real behavior, determine which bookings leave the pool.
  4. Dispatch happens. Bookings are batched into trips, trips are matched to drivers by an optimization solver, and a recycling step decides, for each trip, whether to send it now, hold it, or split it back into bookings and try again later.
  5. Post-dispatch cancellations are applied. These include passenger and driver cancellations.
  6. Outcomes are recorded. Every booking, trip, and driver outcome is written to the run outputs.

Crucially, this is a closed loop: today’s dispatch decisions change where drivers end up, which changes what’s possible next tick. That feedback produces second-order effects that a static replay, one that scores historical decisions without updating future supply, cannot show. Supply may dry up in a hot zone, or one bad dispatch rule may cascade into a wave of cancellations.

Marketplace simulation tick loop from demand through dispatch to recorded outcomes
Figure 1. Simulator sequence.

That loop is only the mechanism. Making it a lab takes three further properties: configurability; reproducibility and auditability; and reliable comparison.

1. Interchangeable stages. The major dispatch stages are pluggable. Batching, allocation, cancellation, routing, driver movement, and recycling are each exposed through an interface with interchangeable implementations selected by configuration. This turns a fixed pipeline into an experimental platform: the component under study can be swapped while everything around it stays unchanged. The same mechanism determines which marketplace is being simulated. Ride-hailing, food delivery, parcel delivery, or all three sharing one driver pool are configuration choices within the same codebase, not separate simulators.

2. Self-contained runs. Each experiment is a small, portable package containing its configuration files, input data, source revision, metrics report, and detailed outputs. Anyone holding that package can recreate the setup, whether a reviewer, teammate, agent, or the original author six months later. The report identifies the configured components and headline metrics, while the supplementary per-booking, per-trip, and per-driver outputs let reviewers recompute those metrics and investigate unexpected outcomes without relying on separate notes or systems.

3. Experiments inside the simulation. Production marketplaces have experimentation platforms, so our simulator ships with one too. Experimentable fields can define multiple treatment arms assigned through the same time-based switchbacks, spatial splits, or cell-and-hour schemes used in production. Each booking records its resolved arm for traceability. Running these A/B/n comparisons together exposes every arm to the same demand, drivers, marketplace dynamics, noise, and biases. By reproducing both the conditions and assignment design of a marketplace experiment, simulated insights are more likely to translate into effects observed in production.

Automating the research cycle

Agents do two jobs here, and both need the same environment: searching for strategies that beat the incumbent, and keeping the simulator sufficiently faithful for that search to mean something.

Searching for better strategies

We connect agents to that interface through autoresearch, an automated loop that runs the research cycle end to end. An agent proposes a change, implements it in source code, tests it, and acts on the verdict. A candidate survives only if it improves the target metric and passes every required check; otherwise it is rolled back and the loop continues. The objective may be a marketplace outcome such as orders-throughput or a software measure such as execution time. What matters is a repeatable command-and-metric contract that both humans and agents can use.

Automated research loop connecting an agent to the simulator and evaluation checks

The first failure mode we encountered was specification gaming. Given a target and a loophole, an agent may find the shortest path to the number rather than the improvement we intended. Ours discovered that changing fields used by pre-dispatch cancellation could reduce cancellations and raise completion without improving a single dispatch decision. The metric moved; nothing real had improved. Rather than relying on instructions alone, we built four safeguards into the environment:

  • Core data is immutable. Modules under test may change only the state exposed by their interfaces. An attempt to manipulate protected fields fails to compile instead of producing a misleading result.
  • The verdict is computed, not judged. The simulator computes its own metrics, which the agent can read but not redefine. Executable checks apply calibrated tolerance bands, verify equivalent outputs where required, and run the tests. The agent never judges its own work.
  • Changes are scoped, checked, and reversible. Each experiment runs on its own branch within a declared writable scope, which is checked against the resulting diff. The grading machinery remains outside that scope, regressions cost only a discarded branch, and every surviving change is bounded enough to review in full.
  • Changes are checked at two levels. It first rejects changes that violate type contracts, ownership rules, interface boundaries, or concurrency requirements. The evaluation harness then applies checks suited to the objective: performance work must preserve expected outputs, while strategy experiments must satisfy domain constraints and outcome thresholds.

The speed that matters is the speed of an adaptive research loop, not simulation runtime alone: each iteration uses previous results to propose the next change, then implements, compiles, runs, scores, and decides whether to keep it. The simulator returns a verdict in milliseconds for benchmarks or tens of minutes for a full-day replay, while autoresearch carries each verdict into the next attempt without human handoffs. This turns an overnight run into a connected sequence of evidence-driven experiments, with the safeguards above ensuring that faster iteration compounds reliable evidence rather than mistakes.

One outcome is the familiar purpose of simulation: discovering better marketplace strategies. When asked to explore trip-recycling policy, the loop turned an emerging human intuition into a concrete rule: hold batched orders only as long as service-level deadlines permit, maximizing the chance that another nearby order joins the trip. Because the result is human-readable code, engineers can review, audit, and deploy it like any other pull request.

The same machinery also improves the simulator itself. Over several nights of unattended operation, we aimed the agents at two hot paths in its dispatch logic: the checks that decide which drivers are eligible for a trip, and the pass that adjusts the cost of every driver-and-trip pairing before the solver chooses an assignment. Across roughly 150 logged experiments, three out of four attempts failed to build or were rejected by later gates. The eligibility checks ended up about 10 times faster, and the cost pass about 24 times faster. Because these paths run repeatedly in every replay, improvements compound across subsequent research. Faster and more efficient runs enable testing of more ideas across more markets and dates, repeat runs to separate signal from noise, and validate results more rigorously. Shorter runs also tighten the feedback loop from verdict to next proposal, so improving the instrument accelerates and strengthens every search performed with it.

Execution time is only one possible objective: pointing the same machinery at orders-throughput changes the research question, not the loop itself. Regardless, whatever the objective, the result is only as trustworthy as the simulator behind it. An agent can optimize only the world it is given; if that world has drifted from production, a faster loop will merely produce misleading answers sooner.

Keeping the simulator faithful

A simulator is a claim about the world: that this is how the system behaves and that this is how the people within it respond. Both halves of that claim decay. Production logic changes continuously, while models of passenger and driver behavior grow stale as new product features reshape how people act. A simulator that has drifted from the world produces misleading insights and innovations that fail in production. So the second job we give agents is to keep the simulator faithful.

The system half is a translation problem, and coding agents are particularly good at it. An agent can read a component’s production implementation and implement equivalent logic behind the simulator’s corresponding interface, translating directly from source code rather than from a written description that may already be stale. The result remains a candidate until parity checks show that it matches the production behavior being modeled.

The behavioral half cannot be translated, because no source file tells us how a passenger will behave. It must be inferred from what people actually do. Here, we redirect the same autoresearch loop from dispatch-code optimization to behavioral modeling. It proposes, fits, and scores candidate models, exploring which features predict cancellation and how their effects should be represented. The test is not only how closely a model explains its training data, but also how well it generalizes to held-out data. A model that fits last Tuesday perfectly and next Tuesday poorly does not make the simulator more faithful; it makes the lab more confidently wrong.

The pieces are simple: a realistic marketplace model, swappable parts, repeatable experiments, honest scoring, and an easy way to undo failures. Together, they give agents a safe place to test and improve ideas. Everyone is racing to give AI a bigger brain. We got further by giving it a better lab: a place where it can try a thousand ideas, be wrong cheaply, and receive an honest verdict.

Join us

Grab is Southeast Asia’s leading superapp, serving over 900 cities across eight countries (Cambodia, Indonesia, Malaysia, Myanmar, the Philippines, Singapore, Thailand, and Vietnam). Through a single platform, millions of users access mobility, delivery, and digital financial services, including ride-hailing, food delivery, payments, lending, and digital banking via GXS Bank and GXBank. Founded in 2012, Grab’s mission is to drive Southeast Asia forward by creating economic empowerment for everyone while delivering sustainable financial performance and positive social impact.

Powered by technology and driven by heart, our mission is to drive Southeast Asia forward by creating economic empowerment for everyone. If this mission speaks to you, join our team today!

DispatchGym: Grab’s reinforcement learning research framework

Post Syndicated from Grab Tech original https://engineering.grab.com/techblog_-dispatchgym

Introduction

DispatchGym is a research framework designed to facilitate Reinforcement Learning (RL) studies and applications for the dispatch system, which matches bookings with drivers. The primary goal is to empower data scientists with a tool that allows them to independently develop and test RL-related concepts for dispatching systems. It accelerates research by providing a suite of modules that include a reinforcement learning algorithm, a dispatching process simulation, and an interface connecting the two through the Gymnasium API.

To ensure efficient and cost-effective RL research without compromising on quality, DispatchGym aims to be both comprehensive and accessible. Anyone with basic RL knowledge and Python programming skills can use it to explore new ideas in RL and dispatch system logic.

This article walks you through the principles behind DispatchGym, how these principles effectively and efficiently empower impactful research, and how it can be applied to solve real world problems.

The challenge with RL

Although RL methods can be applied to a wide variety of problems that can be formulated as a Markov Decision Process (MDP), designing an effective RL-based solution is not a trivial task. The primary challenges stem from two key components: the reward function and the lever.

In RL, the reward function represents the objective we aim to maximize. At first glance, it might seem straightforward to plug in any metric, such as the company’s profit or the number of completed bookings per day. However, these metrics are not always sensitive to the lever that RL can manipulate, or the lever itself may not significantly influence the objective. For example, consider a setup where we aim to maximize the daily number of completed bookings by adjusting the maximum number of candidate drivers considered to each booking. Beyond a minimal threshold (e.g., one driver), further increasing this limit provides negligible benefits. As a result, RL struggles to determine whether setting this limit to 11 or 15 would result in higher rewards.

In summary, when a lever exerts weak influence on a reward function, the RL setup becomes ineffective. Therefore, we should strive to select a lever that strongly influences the reward function and define a reward function that is both sensitive to manipulations of that lever and aligned with our overall goal. Note that the reward function does not have to be identical to our ultimate objective; it merely needs to be highly correlated with it.

Figure 1. Illustration of weak lever influence on a reward function.

Empowering research with DispatchGym

The primary application of DispatchGym is to accelerate and broaden cost-effective research and impactful RL applications for Grab’s dispatching system. A system which is responsible for assigning a driver to each booking. To achieve this, DispatchGym must have the following characteristics:

  • Reliable
    The simulation component should be accurate enough to capture essential behaviors strongly linked to the metrics of interest, without necessarily modeling everything else. While it’s beneficial if the simulation can do more than the specific use case (e.g., simulating both batching and allocation when only allocation is needed), it is not strictly required.

  • Cost-effective
    Updating all of DispatchGym’s components should require minimal monetary and labor costs to enable rapid iteration. This includes keeping the simulation component aligned with real system behaviors, incorporating the latest technologies in the optimization component, and maintaining seamless integration between the simulation and optimization components.

  • Empowering
    It should be as easy as possible for data scientists and engineers to modify any DispatchGym component and then run experiments. This flexibility is crucial because new research typically requires adjustments to both the simulation and optimization components. By granting users the freedom to adapt DispatchGym, the framework fosters continuous innovation.

Research-friendly simulated environment

The simulation component of DispatchGym, or the “simulated environment,” is designed with reliability, cost-effectiveness, and user empowerment in mind. It models the full dispatching process, from booking creation and driver dispatch to driver movement and booking completion. While this environment may not be perfectly accurate in absolute terms (there can be differences between real and simulated metric values), it emphasizes directional accuracy. This means that the metric trends (up or down) in the simulation closely match real-world behavior. This focus on directional accuracy is crucial because most research involves sim-to-sim comparisons, where shifts in metrics are the most important. Verifying directional accuracy is also simpler and more practical for evaluating simulation performance. For instance, we can test various supply-demand imbalance scenarios and check whether a supply-rich situation indeed fulfills more bookings, and vice versa.

Figure 2. Simulated processes.

The simulated environment’s cost-effectiveness and empowerment features come from a modular architecture and Python, a research-friendly programming language. The modular design offers a gentle learning curve, allowing users to easily navigate and make necessary changes in the codebase. Meanwhile, Python is selected to lower the entry barrier for adopting DispatchGym. To mitigate Python’s runtime overhead, DispatchGym leverages Numba to significantly speed up simulation execution.

DispatchGym in action

Data scientists use DispatchGym by modifying a local copy of the codebase to implement their ideas. They then upload the updated codebase to an internal infrastructure using a single CLI command, which spawns a Spark job to run the DispatchGym program. This setup grants complete flexibility over the simulation and optimization components without requiring users to manage the underlying infrastructure.

Figure 3. Data scientist interactions with DispatchGym.

Applying RL approach for dispatch

Amongst its many uses, DispatchGym was applied in building an effective contextual bandit strategy for the auto-adaptive tuning of dispatch-related hyperparameters. Its flexibility allowed us to experiment with various contextual bandit model variants, including linear bandits, neural-linear bandits, and Gaussian-process bandits, as well as multiple action sampling strategies, such as epsilon-greedy, Thompson sampling, SquareCB, and FastCB. These capabilities accelerated our progress in determining the best combination of levers, reward functions, and contextual bandits for improved fulfilment efficiency and reliability.

Conclusion

DispatchGym provides us a framework that equips data scientists with everything they need to develop and test RL solutions for dispatch systems. By integrating an RL optimization approach and a realistic dispatch simulation using a Gymnasium API, it enables rapid exploration and iteration of RL applications with just basic RL knowledge and Python programming language.

A major hurdle in applying RL to dispatch problems modeled as MDP is ensuring that the reward function aligns with ultimate business goals and is sensitive to the lever under control. If the lever (e.g., tweaking driver count) does not meaningfully influence the reward, the RL approach falters. DispatchGym addresses this by making it easy for data scientists to determine the most effective combinations of levers, reward functions, and RL approaches, ultimately driving positive business impact.

DispatchGym’s architecture focuses on reliability, cost-effectiveness, and user empowerment. Its simulation is designed to capture critical metrics and reflect real-world trends (directional accuracy), while its Python-based modular design enhanced by Numba enables easy prototyping. Researchers can adjust the environment locally before deploying changes seamlessly via a command-line interface, avoiding infrastructure overhead. These design decisions and capabilities empower data scientists to refine contextual bandit approaches for optimizing dispatch hyperparameters and explore innovative RL applications in the dispatch process.

We would like to thank Chongyu Zhou, Guowei Wong, and Roman Kotelnikov for their collaboration in developing the RL-based optimizer.

Join us

Grab is a leading superapp in Southeast Asia, operating across the deliveries, mobility and digital financial services sectors. Serving over 800 cities in eight Southeast Asian countries, Grab enables millions of people everyday to order food or groceries, send packages, hail a ride or taxi, pay for online purchases or access services such as lending and insurance, all through a single app. Grab was founded in 2012 with the mission to drive Southeast Asia forward by creating economic empowerment for everyone. Grab strives to serve a triple bottom line – we aim to simultaneously deliver financial performance for our shareholders and have a positive social impact, which includes economic empowerment for millions of people in the region, while mitigating our environmental footprint.

Powered by technology and driven by heart, our mission is to drive Southeast Asia forward by creating economic empowerment for everyone. If this mission speaks to you, join our team today!