Heritage
Control theory and reinforcement learning in finance, traced backwards.
Five strands braid through this history: optimal control and dynamic programming; the portfolio and hedging problems of mathematical finance; equilibrium models that are secretly control problems; dealer quoting and order execution; and the trial-and-error line that became reinforcement learning.
The strands cross more often than the textbooks admit. Rational expectations was born inside inventory control (Muth 1961). The earliest published temporal-difference rule appeared in a control journal (Witten 1977). Mean–variance was in print twice before Markowitz — once by his own thesis adviser. And the 1958 grain-storage equilibrium is now computed by reinforcement learning. Citation links below were chased backwards from the modern papers to their sources; each entry states what it built on.
control portfolio equilibrium dealers execution rl industry jpm — the last marks the thread through J.P. Morgan, recorded here as first-hand testimony; pull requests from others in industry are most welcome.
Deep background, 1654–1892
The problem-of-points correspondence. Mathematical expectation invented to divide the stakes of an interrupted game — the first formal valuation of an uncertain claim. Pascal’s wager (posthumous, 1670) extends it to decision under uncertainty.
De Ratiociniis in Ludo Aleae. The first published treatise on expectation, systematizing what Huygens learned of the Pascal–Fermat exchange; the standard probability text for half a century, reprinted as Part I of Ars Conjectandi.
The brachistochrone challenge (Acta Eruditorum). Solved by Johann, Jacob, Leibniz, l’Hôpital and Newton, it founds the calculus of variations (Euler 1744, Lagrange) — the variational root optimal control still cites (Sussmann & Willems, “300 Years of Optimal Control”). The Bernoulli family anchors both branches of this page.
Jacob’s Ars Conjectandi (posthumous) delivers the law of large numbers; the same year, Nicolas poses the St. Petersburg game in a letter to Montmort — the problem that will break raw expectation as a decision criterion.
A letter to Nicolas Bernoulli proposes concave (square-root) valuation of payoffs to defuse St. Petersburg — utility before the name. Daniel Bernoulli quotes the letter, in French, at the end of his 1738 paper.
Specimen theoriae novae de mensura sortis. Answers the problem Nicolas posed in 1713: gamblers maximize expected log wealth. The utility function returns as the growth-optimal criterion (Kelly, Breiman) and as the borderline CRRA case Merton flags in 1969.
An Essay towards Solving a Problem in the Doctrine of Chances, generalized by Laplace (1774). Inverse probability — belief updated by evidence. The forward link is state estimation: the Kalman filter is recursive Bayesian inference, made explicit by Ho & Lee (1964).
On a General Method in Dynamics (Jacobi’s generalization, 1837). Mechanics recast as a PDE for a characteristic function — the HJ of HJB. The route ran through Carathéodory’s 1935 “royal road”; Bellman arrived independently, and it was Kalman who joined the names.
On Governors (Proc. Royal Society). The first stability analysis of a feedback system — linearize Watt’s governor, demand roots with negative real parts. Neglected for eighty years until Wiener’s Cybernetics canonized it as the founding paper of control.
The Mathematical Theory of Banking. A bank’s cash reserve against stochastic withdrawals — the newsvendor problem avant la lettre, now recognized as the ancestor of stochastic inventory theory (Arrow–Harris–Marschak reinvented it without citation).
The General Problem of the Stability of Motion. Stability certificates without solving the dynamics; Kalman & Bertram carried the “second method” into control design in 1960 — the bridge into the state-space era.
Feedback, storage and expectation, 1900–1948
Théorie de la spéculation. Prices as Brownian motion, five years before Einstein — the state dynamics that continuous-time control in finance would later steer. Rediscovered via Savage and Samuelson in the 1950s.
Animal Intelligence. The law of effect: actions followed by satisfaction are strengthened. The psychological root of the “reinforcement” in reinforcement learning.
How Many Parts to Make at Once (Factory). The square-root economic order quantity — the founding formula of inventory theory, miscredited to Wilson for seventy-five years.1
Differential-Space. Brownian motion made rigorous as a measure on path space — the Wiener process, the state noise of everything continuous-time on this page, and the integrand of Itô’s calculus.
Black’s negative feedback amplifier (conceived 1927), Nyquist’s Regeneration Theory (1932) and Bode’s gain–phase analysis (1940). The Bell Labs feedback tradition that Wiener’s Cybernetics would generalize and name.
A Mathematical Theory of Saving. Optimal national saving via the calculus of variations — the first economic control problem. Merton and Samuelson both later name their reduced consumption problem the “Phelps–Ramsey problem.”
Zur Theorie der Gesellschaftsspiele. Minimax and backward reasoning through the game tree — ancestor of both expected utility (1944) and Wald’s minimax statistics.
The Economics of Exhaustible Resources. The price of a stored exhaustible asset rises at the rate of interest — an equilibrium derived from the holder’s intertemporal optimization, equilibrium-as-control two decades before Bellman.
On the likelihood that one unknown probability exceeds another (Biometrika). Sample actions from the posterior probability that they are best: the earliest bandit algorithm — and the one running live in Man AHL’s order router eight decades later.
Working’s wheat-futures empirics (1933) show spot–futures spreads pricing the service of storage; Kaldor names the convenience yield (1939); Working states the supply-of-storage theory (AER, 1949). With Brennan (1958), the line that feeds Gustafson and Deaton–Laroque.
Marschak’s Money and the Theory of Assets orders risky positions by means and covariances; de Finetti’s Il problema dei “pieni” (1940) computes a mean–variance frontier for reinsurance. Mean–variance twice before Markowitz — Marschak supervised Markowitz’s dissertation without ever mentioning his own paper, and Markowitz conceded the second in de Finetti Scoops Markowitz (2006).
Kolmogorov’s Interpolation und Extrapolation von stationären zufälligen Folgen and Wiener’s classified 1942 anti-aircraft study (published 1949 as Extrapolation, Interpolation, and Smoothing of Stationary Time Series). Optimal prediction from noisy observation, twice and independently — Wiener’s report acknowledges the parallel. The ancestor of every estimator that later extracts signal from prices.
Theory of Games and Economic Behavior. Axiomatizes expected utility — the objective function every stochastic-control formulation in finance maximizes.
Stochastic Integral (1944) and On Stochastic Differential Equations (Memoirs AMS, 1951). The calculus of controlled diffusions, built on Wiener’s process; Merton’s 1969 bibliography carries it as reference [4]. Doob’s Stochastic Processes (1953) adds the martingale toolkit that measure-change pricing later rests on.
Les réserves et la régulation de l’avenir. Hydroelectric reservoirs managed as sequential decisions under uncertainty — storage control, and a recursion anticipating dynamic programming, a decade before Bellman named it.
Sequential Analysis. Decisions made as data arrive, with a cost of continuing to sample. Ancestor of both Bellman’s dynamic programming and the optimal-stopping problem that American-option pricing lives in.
Cybernetics. Feedback control, communication, and learning unified across animals and machines — the manifesto that made “a controller that learns” a legitimate scientific object.
The dynamic-programming decade, 1949–1965
Bayes and Minimax Solutions of Sequential Decision Problems. Wald’s sequential analysis recast as backward-induction functional equations — the explicit bridge from statistical decision theory to dynamic programming, co-authored by two other principals of this page.
A Stochastic Approximation Method. Root-finding through noise with decaying step sizes. The mathematical engine of temporal-difference and Q-learning: the 1994 convergence proofs rest on it, and its step-size conditions appear verbatim in Q-learning’s.
Games of Pursuit (RAND P-257). Differential games and the Hamilton–Jacobi–Isaacs equation, developed at RAND independently of Bellman down the hall; the memoranda followed in 1954–55, the book in 1965.
Portfolio Selection. The static mean–variance problem. Every dynamic theory below is an answer to “what happens to Markowitz over time?”
Some Aspects of the Sequential Design of Experiments. The multi-armed bandit posed as the canonical exploration–exploitation problem, generalizing Thompson; resolved by dynamic programming in Gittins’ index theorem (1974, 1979).
Applications of Martingale System Theorems. The smallest supermartingale dominating a reward — the Snell envelope — solves optimal stopping; three decades on it is the value process of the American option.
The Transactions Demand for Cash: An Inventory Theoretic Approach. Money demand as a lot-size problem — cash is “its holder’s inventory of the medium of exchange.” With Tobin (1956), the inventory view of money; Baumol credits the 1920s lot-size literature rather than Harris, whose paper was then lost.
Stochastic Games. Value iteration on a discounted dynamic model, four years before Bellman’s book — restricted to one player it is value iteration, and the two-player frame returns in multi-agent RL and market games.
A New Interpretation of Information Rate. A gambler with a private noisy channel maximizes the exponential growth rate of capital, and that rate equals Shannon’s information rate — Bernoulli’s log utility meets information theory. Breiman (1961) proves the log-optimal strategy asymptotically optimal.
Dynamic Programming and A Markovian Decision Process. The principle of optimality, the Bellman equation, and the MDP. Everything on the algorithmic side of this page reduces to solving his equation.
Carryover Levels for Grains (USDA). Competitive commodity storage as a Bellman recursion, with the crucial non-negativity constraint on inventories — the first equilibrium computed as a dynamic program, solved (per Deaton–Laroque) under rational expectations before Muth’s paper existed.
Some Studies in Machine Learning Using the Game of Checkers. An evaluation function updated toward the value of later positions — temporal-difference learning avant la lettre.
Dynamic Programming and Markov Processes. Policy iteration: alternate evaluation and greedy improvement. Modern actor-critic methods are its sampled, incremental descendants.
A New Approach to Linear Filtering and Prediction Problems, with the linear-quadratic regulator alongside — the exactly solvable case of Bellman’s problem and the benchmark learning controllers are measured against. Pontryagin’s maximum principle (English translation 1962) supplies the continuous-time counterpart, developed independently in the USSR; Kalman & Bucy (1961) add the continuous-time filter, and Stratonovich and Zakai carry filtering nonlinear.
Planning Production, Inventories, and Work Force. The Carnegie project: quadratic costs give linear decision rules for production and inventory — LQ control before the name, estimated on factory data. The incubator of the next entry.
Rational Expectations and the Theory of Price Movements. Expectations “are essentially the same as the predictions of the relevant economic theory.” Verified in the text: Muth cites HMMS chapters 2–4 and 19, and his centerpiece example is a commodity market with inventory speculation. Rational expectations was born inside inventory control.
The Accumulation of Risky Capital. Ramsey’s control problem made stochastic and sequential by dynamic programming. The verified, directly cited parent of both 1969 lifetime-portfolio papers.
Rational Theory of Warrant Pricing (Industrial Management Review), with McKean’s appendix A Free Boundary Problem for the Heat Equation. American-option pricing posed and solved as optimal stopping, eight years before Black–Scholes; Samuelson & Merton (1969) put it inside a utility equilibrium — the bridge to 1973.
Discounted Dynamic Programming. The discounted MDP of Bellman and Howard put on rigorous measure-theoretic footing, with Blackwell optimality — the theorem base under every convergence guarantee later claimed in RL.
Control enters finance, 1965–1990
Optimal Control of Markov Processes with Incomplete State Information. The POMDP: the belief state is a sufficient statistic, and dynamic programming runs over it. His adaptive-control program is the control-side twin of RL — Sutton, Barto & Williams later call RL “direct adaptive optimal control.”
A Model of the Demand for Money by Firms. Firm cash under a random walk with fixed transfer costs: do nothing inside a band, jump to a return point on hitting either edge — the two-threshold-with-return-point policy that reappears in commodity storage bands and dealer inventory bounds. Eppen & Fama (1969) — Fama doing inventory control — and Constantinides (1976, 1978) make the bands rigorous.
The Cost of Transacting. Immediacy is a costly service and the bid–ask spread is its price — the founding paper of market microstructure economics, in the reference lists of Stoll (1978) and Grossman–Miller (1988).
Lifetime Portfolio Selection by Dynamic Stochastic Programming and Lifetime Portfolio Selection under Uncertainty: The Continuous-Time Case, back to back in REStat 51(3). Samuelson solves lifetime consumption and portfolio choice by discrete-time DP, building on Phelps and Mossin (1968); Merton brings Itô calculus and the Hamilton–Jacobi–Bellman equation into finance, yielding the Merton fraction.
Optimum Consumption and Portfolio Rules in a Continuous-Time Model. The general many-asset theory; the 1973 ICAPM turns the same machinery into equilibrium asset pricing with intertemporal hedging demands.
Investment Under Uncertainty. Rational-expectations industry equilibrium shown to be the solution of a planner’s dynamic program — the template “equilibrium = value function of a control problem,” built explicitly on Muth.
Stochastic Speculative Price. The price of a storable commodity “determined as the solution to a stochastic-dynamic-programming problem”; proves the competitive speculative equilibrium coincides with the planner’s optimum, in Gustafson’s setting.
Two Financial Analysts Journal pieces stake out the field’s two mechanisms: The Only Game in Town — “Walter Bagehot,” pseudonym of Jack Treynor — has dealers losing to informed traders and recouping from liquidity traders; Smidt has the dealer as an active price-setter managing inventory. The two lines meet in 1985.
The Pricing of Options and Corporate Liabilities and Theory of Rational Option Pricing. Dynamic replication: the delta hedge is a feedback control law that renders the payoff attainable, collapsing pricing to a PDE. Harrison–Kreps (1979) and Harrison–Pliska (1981) then recast it by duality: no-arbitrage means a martingale measure exists.
Optimal Stochastic Linear Systems with Exponential Performance Criteria. Risk-sensitive (LEQG) control: exponential-of-quadratic cost, certainty equivalence fails, and the optimal controller equals a deterministic game. Whittle (1981, 1990) completes the theory; the same exponential criterion returns in Hansen–Sargent robustness and in risk-sensitive asset management.
Market Microstructure. Coins the term; a dealer facing Poisson order arrivals sets quotes to avoid gambler’s ruin — inventory first enters as a viability constraint. The origin point of the quoting strand; Benston & Hagerman (1974) had supplied the OTC spread empirics.
Portfolio Selection with Transactions Costs. Proportional costs added to Merton’s problem and the no-trade cone conjectured; Constantinides (1986) shows investors optimally tolerate wide drift from target.
An Adaptive Optimal Controller for Discrete-Time Markov Environments. Contains the earliest known published temporal-difference rule — in a control journal. Sutton and Barto discovered it only in 1981, while finishing the actor-critic work.
Rules Rather than Discretion. Optimal control of policy fails when the controlled agents solve their own forward-looking problems: the policymaker’s problem carries the private sector’s equilibrium response as a constraint, making discretionary plans time-inconsistent.
The Supply of Dealer Services in Securities Markets. The spread as compensation for a risk-averse dealer’s displacement from his desired portfolio — risk-bearing cost, not market power.
Asset Prices in an Exchange Economy. A representative agent’s Bellman equation prices all assets. Cox, Ingersoll & Ross (1985) supply the continuous-time counterpart: an HJB equation generating the equilibrium term structure.
Dealership Market: Market-Making with Inventory. Builds on Garman: bounded inventory, quotes adjusted dynamically as functions of the inventory state, a preferred inventory position.
Optimal Dealer Pricing Under Transactions and Return Uncertainty. The canonical stochastic-DP dealer: price-sensitive Poisson arrivals (adopted from Garman), return diffusion, risk aversion; inventory shifts quote levels, not spread. The paper the entire modern quoting literature descends from — then a 27-year gap in the line.
Impulse Control and Quasi-Variational Inequalities (with the companion variational-inequalities volume). The PDE machinery for (s, S)-type policies: the value function of an impulse problem characterized as the solution of a QVI. Harrison, Sellke & Taylor (1983) solve the clean Brownian case; Korn (1998) carries the machinery into Merton’s problem.
Neuronlike Adaptive Elements That Can Solve Difficult Learning Control Problems. The actor-critic architecture — a TD critic training a trial-and-error actor — balancing a pole. Develops Klopf’s (1972) hedonistic-neuron program.
Bid, Ask and Transaction Prices in a Specialist Market and Continuous Auctions and Insider Trading. The adverse-selection alternative made formal: Glosten–Milgrom’s spread exists at zero inventory, explicitly naming Ho–Stoll, Amihud–Mendelson and Garman as the contrast; Kyle’s insider solves a dynamic control problem while the market maker filters order flow — control plus filtering inside one equilibrium.
Optimal Portfolio Choice Under Incomplete Information (with Detemple’s equilibrium version the same year). Filtering formally enters portfolio choice — Kalman-filter the drift, then optimize on the estimate — made unavoidable by Merton’s 1980 demonstration that drifts are what markets never let you measure. Lakner (1995, 1998) makes it rigorous.
Liquidity and Market Structure. Demsetz’s immediacy made equilibrium: the number of market makers set by the cost of continuous presence against the value of immediacy to outside investors; among its cited antecedents, Cohen–Maier–Schwartz–Whitcomb’s proof (1981) that transaction costs alone sustain a spread.
Learning to Predict by the Methods of Temporal Differences. TD(λ) defined and named, with convergence in mean for linear features; credits Samuel, Witten, and the Widrow–Hoff rule as antecedents.
Learning from Delayed Rewards. The thesis that married trial-and-error learning to Bellman–Howard dynamic programming, introducing Q-learning. Watkins & Dayan (1992) prove convergence under exactly the Robbins–Monro step-size conditions. Stokey & Lucas’s Recursive Methods (1989), the same year, codifies DP as the language of economic equilibrium.
Portfolio Selection with Transaction Costs. The Magill–Constantinides problem solved rigorously as singular stochastic control: reflect the portfolio at the boundary of the no-trade wedge. Shreve–Soner (1994) redo it with viscosity solutions.
Learning meets markets, 1992–2001
REINFORCE: unbiased Monte-Carlo policy gradients — optimize the policy directly, not through a value function. And TD-Gammon: TD(λ) plus a neural network plus self-play reaches master-level backgammon, the existence proof cited by both the DQN and AlphaGo papers.
On the Behaviour of Commodity Prices (and JPE 1996). The Gustafson–Samuelson storage equilibrium estimated on thirteen commodity series, via the numerical revival by Wright & Williams (1982) and Scheinkman & Schechtman (1983). The DP-generated price function explains spikes and skewness; autocorrelation less well.
On the Convergence of Stochastic Iterative Dynamic Programming Algorithms and Asynchronous Stochastic Approximation and Q-learning. Q-learning and TD proved convergent by embedding them in Robbins–Monro stochastic approximation — the 1951 foundation made load-bearing.
Neuro-Dynamic Programming. RL systematized as approximate dynamic programming for the control and OR communities — the idiom through which it entered financial engineering. Tsitsiklis & Van Roy (1997) delimit convergence with function approximation.
Gauge theory reaches economics and finance: Malaney’s Harvard thesis builds a gauge connection for index numbers (the Malaney–Weinstein connection); Ilinski (1997; book 2001) reads arbitrage as the curvature of a discounting connection; Young (1999) renders FX as a lattice gauge theory. None of it touches dealer markets: as of August 2026 no published work applies gauge or symmetry language to market making — the gap the final entry on this page occupies.
Optimization of Trading Systems and Portfolios, then Moody & Saffell, Learning to Trade via Direct Reinforcement (2001). Recurrent RL learns the trading policy directly, maximizing a differential Sharpe ratio rather than forecasting prices — the founding direct-RL trading papers.
Optimal Control of Execution Costs. Dynamic programming for buying a block over a fixed horizon under linear impact; the risk-neutral answer is the even split. Opens the execution branch.
Risk-Sensitive Dynamic Asset Management (with Fleming & Sheu, 2000–2002, in parallel). Jacobson–Whittle’s exponential criterion arrives in portfolio theory as factor-based allocation trading long-run growth against variance; Björk, Davis & Landén (2010) consolidate the partial-information line.
Optimal Execution of Portfolio Transactions. Risk aversion added to Bertsimas–Lo (cited directly): an efficient frontier of liquidation schedules trading impact cost against volatility risk, with permanent versus temporary impact. Obizhaeva & Wang (working paper 2005) later add order-book resilience.
Equilibrium Forward Curves for Commodities. The storage DP enters mainstream finance: forward curves from a Deaton–Laroque-style equilibrium, with the inventory non-negativity constraint creating an embedded timing option and endogenous convenience yield.
An Electronic Market-Maker (MIT AI Memo). The first reinforcement-learning market maker: policy-gradient and TD agents quoting in a simulated Garman-style order-flow market, balancing profit and inventory. Cited by both Nevmyvaka–Feng–Kearns and Spooner et al.
Valuing American Options by Simulation and Regression Methods for Pricing Complex American-Style Options. Approximate dynamic programming arrives in finance proper: regress continuation values along simulated paths — approximate value iteration for the Wald-lineage stopping problem. Rogers (2002) supplies the martingale dual. Hansen & Sargent (2001, book 2008) meanwhile import H∞ robust control into economics.
Convergence, 2003–2013
Logarithmic Market Scoring Rules (working paper 2002; with “Combinatorial Information Market Design,” 2003). An automated market maker as a scoring rule: log-sum-exp cost function, bounded worst-case loss — quoting without a quoting desk. Chen & Pennock (2007) recast it as inventory-utility; Othman et al. (2010) make its liquidity adaptive.
Competitive Algorithms for VWAP and Limit Order Trading. Worst-case guarantees from online algorithms — a genuinely separate computer-science tradition (its reference list contains no Bertsimas–Lo, no Almgren–Chriss), merged with the control tradition only by the next entry.
Reinforcement Learning for Optimized Trade Execution. The first large-scale empirical RL for execution — Q-learning over millisecond order-book states on NASDAQ data — and the verified bridge: it cites Bertsimas–Lo, Almgren–Chriss, the VWAP line, and Chan–Shelton. The lead in Halperin’s Quora answer that seeded this page; Nevmyvaka went on to head machine-learning research at Morgan Stanley.
Nash Certainty Equivalence (2006) and Mean Field Games (2007), independently. A continuum of agents each solving a control problem against the population distribution; equilibrium is a coupled backward-HJB / forward-Kolmogorov system — equilibrium-as-control made literal. Carmona & Delarue (2018) give the probabilistic codification.
High-Frequency Trading in a Limit Order Book. Ho–Stoll 1981 revived for electronic books — the paper says so itself (“closely related to a paper by Ho and Stoll”) — with exponential-utility reservation prices and fill intensities λ(δ) = Ae−kδ. Ends the 27-year gap in the quoting line.
Dealing with the Inventory Risk. Inventory bounds turn the Avellaneda–Stoikov HJB into linear ODEs: the first verification theorem and closed-form quotes as functions of inventory. Guilbaud & Pham (2013) add market orders; Cartea, Jaimungal & Penalva’s 2015 textbook makes it the standard curriculum.
Dynamic Trading with Predictable Returns and Transaction Costs. The synthesis in closed form: Markowitz’s statics, Merton’s dynamics, and transaction costs resolved with the LQ regulator — aim in front of the target.
Playing Atari with Deep Reinforcement Learning (Nature version 2015; AlphaGo 2016). Watkins’ Q-learning stabilized with deep networks, experience replay and target networks — citing Watkins–Dayan and Tesauro. The enabling technology for everything below; DDPG (2015) and PPO (2017) supply the continuous-action line finance mostly uses.
Into production, 2014–2026
Theoretically motivated width and skew for OTC trading, and the discovery of a key gauge symmetry making control and RL feasible — in production at J.P. Morgan, begun on arrival in 2013 and live the following year. The theory circulated as Cotton & Papanicolaou, Trading Illiquid Goods (working paper; presented at NYU, 2018); the arXiv record of the mid-2010s work is the final entry on this page. With the assistance of Nick West, Erik Clossen and Pascal Tomecek.
J.P. Morgan’s deep-RL execution engine, reported July 2017 as live in European equities since Q1, trained on billions of historic and simulated transactions.
A learning order router in live production, per Risk.net’s fund-of-the-year citation: trades allocated across internal algos, dealer algos and the human desk by Thompson sampling — a 1933 bandit algorithm running money. Bloomberg had reported Man Group’s machine-learning models live across $12.7bn of funds that July.
Machine Learning for Trading: Q-learning with a mean–variance reward discovers statistical arbitrage under costs — the paper that made RL respectable in the practitioner quant literature. And QLBS: option price and hedge emerge as the optimal Q-function of a risk-adjusted MDP — Black–Scholes done by Q-learning. Kolm & Ritter’s 2019 survey maps the field onto the RL formalism.
Deep Hedging (Quantitative Finance 2019). A derivatives book hedged by deep RL under transaction costs and risk limits, sidestepping complete-market pricing — written inside J.P. Morgan’s equities QR. In production hedging vanilla index flow from 2018; by 2022, per Risk.net, 70% of the bank’s Euro Stoxx flow options were quoted and hedged by machine, with Deep Bellman Hedging (2022) the actor-critic sequel. Buehler left for the market maker XTX in 2022.
Idiosyncrasies and Challenges of Data Driven Learning in Electronic Trading. The bank’s own candid account of RL in production trading: thousands of micro-decisions, reward-shaping difficulties, regulatory constraints — “we backtest, therefore we overfit.” The same year, Manuela Veloso founds J.P. Morgan AI Research, whose ABIDES simulator (2019) and multi-agent RL dealer markets (2019) carry the simulation line forward.
Market Making via Reinforcement Learning. TD agents with an asymmetrically dampened inventory reward recover skewing behavior on high-frequency order-book data — citing both lineages: Ho–Stoll and Avellaneda–Stoikov on one side, Chan–Shelton and Nevmyvaka–Feng–Kearns on the other. Cardaliaguet & Lehalle’s mean field game of controls lands MFGs in microstructure the same year.
Deep Reinforcement Learning for Market Making in Corporate Bonds. Actor-critic networks replace finite-difference HJB solvers for multi-asset OTC quoting — the point where the Ho–Stoll control lineage and the RL lineage formally merge, on exactly the asset class of the 2014 entry.
Learning Mean-Field Games. MFG equilibrium computed by model-free Q-learning, with guarantees; fictitious play (Perrin et al. 2020) and scalable deep RL (Laurière et al. 2022) follow. With Gomes & Saúde’s price-formation MFG (2021) — the price as a Lagrange multiplier on flow balance for a stored commodity — the loop closes: Gustafson’s 1958 storage equilibrium, solved by reinforcement learning.
Aiden, built with Borealis AI — whose RL program RBC launched in 2017 with Richard Sutton as head academic advisor — goes live: a client-facing deep-RL VWAP execution agent, the clearest non-JPM production deep-RL claim on record, followed by Aiden Arrival (2022) and a deep-hedging stress-testing pilot (2021). Beyond the banks, public production claims stay rare: the proprietary market makers publish almost nothing, and XTX frames its billion-euro ML build-out as supervised forecasting.
Dixon, Halperin & Bilokon’s Machine Learning in Finance (Springer), the first graduate text with a full part on RL; and FinRL, the open-source library that became the field’s dominant experimental ecosystem.
Uniswap’s x·y = k deployed at scale (v2 whitepaper, 2020; formal analysis Angeris et al., 2019–21): market making reduced to a fixed convex curve — no inventory control at all, the degenerate opposite of Ho–Stoll. Milionis, Moallemi, Roughgarden & Zhang (2022) price the consequence — loss-versus-rebalancing, the CFMM’s Glosten–Milgrom adverse-selection bill — and ZeroSwap (2023) starts handing the curve learned, inventory-style control back.
Recent Advances in Reinforcement Learning in Finance (Mathematical Finance 33(3)). The standard mathematical survey, connecting RL back to the stochastic-control formulations of execution, market making and portfolio choice; the Annual Review of Statistics survey (2025) closes the period, naming explainability and robustness as the barriers to adoption.
On a Simple Relationship Between Order Imbalance, Skew and Width in Over-The-Counter Trading. The imbalance–skew–width symmetry isolated from the mid-2010s market-making work — this project. The dealer strand of this page, begun with Garman’s viability constraint, restated as a symmetry.
1 Harris published the square-root formula in Factory, The Magazine of Management in 1913. In 1931 F. E. Raymond’s book — the first on the subject — garbled the citation, making the original untraceable; R. H. Wilson’s 1934 Harvard Business Review article then popularized the result, which operations research thereafter called the “Wilson formula.” The original was rediscovered only in 1988, and Erlenkotter’s “Ford Whitman Harris and the Economic Order Quantity Model” (Operations Research 38(6), 1990) restored priority. ↩
Citation links were verified against primary sources where linked; the companion literature map covers the inventory-and-pricing control strand in depth.