Heritage

Control theory and reinforcement learning in finance, traced backwards.

Five strands braid through this history: optimal control and dynamic programming; the portfolio and hedging problems of mathematical finance; equilibrium models that are secretly control problems; dealer quoting and order execution; and the trial-and-error line that became reinforcement learning.

The strands cross more often than the textbooks admit. Rational expectations was born inside inventory control (Muth 1961). The earliest published temporal-difference rule appeared in a control journal (Witten 1977). Mean–variance was in print twice before Markowitz — once by his own thesis adviser. And the 1958 grain-storage equilibrium is now computed by reinforcement learning. Citation links below were chased backwards from the modern papers to their sources; each entry states what it built on.

control portfolio equilibrium dealers execution rl industry jpm — the last marks the thread through J.P. Morgan, recorded here as first-hand testimony; pull requests from others in industry are most welcome.

Deep background, 1654–1892

1654·Pascal & Fermatportfolio

The problem-of-points correspondence. Mathematical expectation invented to divide the stakes of an interrupted game — the first formal valuation of an uncertain claim. Pascal’s wager (posthumous, 1670) extends it to decision under uncertainty.

1657·Christiaan Huygensportfolio

De Ratiociniis in Ludo Aleae. The first published treatise on expectation, systematizing what Huygens learned of the Pascal–Fermat exchange; the standard probability text for half a century, reprinted as Part I of Ars Conjectandi.

1696·Johann Bernoullicontrol

The brachistochrone challenge (Acta Eruditorum). Solved by Johann, Jacob, Leibniz, l’Hôpital and Newton, it founds the calculus of variations (Euler 1744, Lagrange) — the variational root optimal control still cites (Sussmann & Willems, “300 Years of Optimal Control”). The Bernoulli family anchors both branches of this page.

1713·Jacob & Nicolas Bernoulliportfolio

Jacob’s Ars Conjectandi (posthumous) delivers the law of large numbers; the same year, Nicolas poses the St. Petersburg game in a letter to Montmort — the problem that will break raw expectation as a decision criterion.

1728·Gabriel Cramerportfolio

A letter to Nicolas Bernoulli proposes concave (square-root) valuation of payoffs to defuse St. Petersburg — utility before the name. Daniel Bernoulli quotes the letter, in French, at the end of his 1738 paper.

1738·Daniel Bernoulliportfolio

Specimen theoriae novae de mensura sortis. Answers the problem Nicolas posed in 1713: gamblers maximize expected log wealth. The utility function returns as the growth-optimal criterion (Kelly, Breiman) and as the borderline CRRA case Merton flags in 1969.

1763·Bayes & Laplacecontrol

An Essay towards Solving a Problem in the Doctrine of Chances, generalized by Laplace (1774). Inverse probability — belief updated by evidence. The forward link is state estimation: the Kalman filter is recursive Bayesian inference, made explicit by Ho & Lee (1964).

1834·Hamilton & Jacobicontrol

On a General Method in Dynamics (Jacobi’s generalization, 1837). Mechanics recast as a PDE for a characteristic function — the HJ of HJB. The route ran through Carathéodory’s 1935 “royal road”; Bellman arrived independently, and it was Kalman who joined the names.

1868·James Clerk Maxwellcontrol

On Governors (Proc. Royal Society). The first stability analysis of a feedback system — linearize Watt’s governor, demand roots with negative real parts. Neglected for eighty years until Wiener’s Cybernetics canonized it as the founding paper of control.

1888·Francis Ysidro Edgeworthcontrol

The Mathematical Theory of Banking. A bank’s cash reserve against stochastic withdrawals — the newsvendor problem avant la lettre, now recognized as the ancestor of stochastic inventory theory (Arrow–Harris–Marschak reinvented it without citation).

1892·Aleksandr Lyapunovcontrol

The General Problem of the Stability of Motion. Stability certificates without solving the dynamics; Kalman & Bertram carried the “second method” into control design in 1960 — the bridge into the state-space era.

Feedback, storage and expectation, 1900–1948

1900·Louis Bachelierportfolio

Théorie de la spéculation. Prices as Brownian motion, five years before Einstein — the state dynamics that continuous-time control in finance would later steer. Rediscovered via Savage and Samuelson in the 1950s.

1911·Edward Thorndikerl

Animal Intelligence. The law of effect: actions followed by satisfaction are strengthened. The psychological root of the “reinforcement” in reinforcement learning.

1913·Ford Whitman Harriscontrol

How Many Parts to Make at Once (Factory). The square-root economic order quantity — the founding formula of inventory theory, miscredited to Wilson for seventy-five years.1

1923·Norbert Wienerportfolio

Differential-Space. Brownian motion made rigorous as a measure on path space — the Wiener process, the state noise of everything continuous-time on this page, and the integrand of Itô’s calculus.

1927·Black, Nyquist & Bodecontrol

Black’s negative feedback amplifier (conceived 1927), Nyquist’s Regeneration Theory (1932) and Bode’s gain–phase analysis (1940). The Bell Labs feedback tradition that Wiener’s Cybernetics would generalize and name.

1928·Frank Ramseyportfolio

A Mathematical Theory of Saving. Optimal national saving via the calculus of variations — the first economic control problem. Merton and Samuelson both later name their reduced consumption problem the “Phelps–Ramsey problem.”

1928·John von Neumannequilibrium

Zur Theorie der Gesellschaftsspiele. Minimax and backward reasoning through the game tree — ancestor of both expected utility (1944) and Wald’s minimax statistics.

1931·Harold Hotellingequilibrium

The Economics of Exhaustible Resources. The price of a stored exhaustible asset rises at the rate of interest — an equilibrium derived from the holder’s intertemporal optimization, equilibrium-as-control two decades before Bellman.

1933·William R. Thompsonrl

On the likelihood that one unknown probability exceeds another (Biometrika). Sample actions from the posterior probability that they are best: the earliest bandit algorithm — and the one running live in Man AHL’s order router eight decades later.

1933·Working & Kaldorequilibrium

Working’s wheat-futures empirics (1933) show spot–futures spreads pricing the service of storage; Kaldor names the convenience yield (1939); Working states the supply-of-storage theory (AER, 1949). With Brennan (1958), the line that feeds Gustafson and Deaton–Laroque.

1938·Marschak & de Finettiportfolio

Marschak’s Money and the Theory of Assets orders risky positions by means and covariances; de Finetti’s Il problema dei “pieni” (1940) computes a mean–variance frontier for reinsurance. Mean–variance twice before Markowitz — Marschak supervised Markowitz’s dissertation without ever mentioning his own paper, and Markowitz conceded the second in de Finetti Scoops Markowitz (2006).

1941·Kolmogorov & Wienercontrol

Kolmogorov’s Interpolation und Extrapolation von stationären zufälligen Folgen and Wiener’s classified 1942 anti-aircraft study (published 1949 as Extrapolation, Interpolation, and Smoothing of Stationary Time Series). Optimal prediction from noisy observation, twice and independently — Wiener’s report acknowledges the parallel. The ancestor of every estimator that later extracts signal from prices.

1944·von Neumann & Morgensternportfolio

Theory of Games and Economic Behavior. Axiomatizes expected utility — the objective function every stochastic-control formulation in finance maximizes.

1944·Kiyosi Itôportfolio

Stochastic Integral (1944) and On Stochastic Differential Equations (Memoirs AMS, 1951). The calculus of controlled diffusions, built on Wiener’s process; Merton’s 1969 bibliography carries it as reference [4]. Doob’s Stochastic Processes (1953) adds the martingale toolkit that measure-change pricing later rests on.

1946·Pierre Massécontrol

Les réserves et la régulation de l’avenir. Hydroelectric reservoirs managed as sequential decisions under uncertainty — storage control, and a recursion anticipating dynamic programming, a decade before Bellman named it.

1947·Abraham Waldcontrol

Sequential Analysis. Decisions made as data arrive, with a cost of continuing to sample. Ancestor of both Bellman’s dynamic programming and the optimal-stopping problem that American-option pricing lives in.

1948·Norbert Wienercontrol

Cybernetics. Feedback control, communication, and learning unified across animals and machines — the manifesto that made “a controller that learns” a legitimate scientific object.

The dynamic-programming decade, 1949–1965

1949·Arrow, Blackwell & Girshickcontrol

Bayes and Minimax Solutions of Sequential Decision Problems. Wald’s sequential analysis recast as backward-induction functional equations — the explicit bridge from statistical decision theory to dynamic programming, co-authored by two other principals of this page.

1951·Robbins & Monrorl

A Stochastic Approximation Method. Root-finding through noise with decaying step sizes. The mathematical engine of temporal-difference and Q-learning: the 1994 convergence proofs rest on it, and its step-size conditions appear verbatim in Q-learning’s.

1951·Rufus Isaacscontrol

Games of Pursuit (RAND P-257). Differential games and the Hamilton–Jacobi–Isaacs equation, developed at RAND independently of Bellman down the hall; the memoranda followed in 1954–55, the book in 1965.

1952·Harry Markowitzportfolio

Portfolio Selection. The static mean–variance problem. Every dynamic theory below is an answer to “what happens to Markowitz over time?”

1952·Herbert Robbinsrl

Some Aspects of the Sequential Design of Experiments. The multi-armed bandit posed as the canonical exploration–exploitation problem, generalizing Thompson; resolved by dynamic programming in Gittins’ index theorem (1974, 1979).

1952·J. L. Snellportfolio

Applications of Martingale System Theorems. The smallest supermartingale dominating a reward — the Snell envelope — solves optimal stopping; three decades on it is the value process of the American option.

1952·William Baumolcontrol

The Transactions Demand for Cash: An Inventory Theoretic Approach. Money demand as a lot-size problem — cash is “its holder’s inventory of the medium of exchange.” With Tobin (1956), the inventory view of money; Baumol credits the 1920s lot-size literature rather than Harris, whose paper was then lost.

1953·Lloyd Shapleyrl

Stochastic Games. Value iteration on a discounted dynamic model, four years before Bellman’s book — restricted to one player it is value iteration, and the two-player frame returns in multi-agent RL and market games.

1956·John Kellyportfolio

A New Interpretation of Information Rate. A gambler with a private noisy channel maximizes the exponential growth rate of capital, and that rate equals Shannon’s information rate — Bernoulli’s log utility meets information theory. Breiman (1961) proves the log-optimal strategy asymptotically optimal.

1957·Richard Bellmancontrol

Dynamic Programming and A Markovian Decision Process. The principle of optimality, the Bellman equation, and the MDP. Everything on the algorithmic side of this page reduces to solving his equation.

1958·Robert Gustafsonequilibrium

Carryover Levels for Grains (USDA). Competitive commodity storage as a Bellman recursion, with the crucial non-negativity constraint on inventories — the first equilibrium computed as a dynamic program, solved (per Deaton–Laroque) under rational expectations before Muth’s paper existed.

1959·Arthur Samuelrl

Some Studies in Machine Learning Using the Game of Checkers. An evaluation function updated toward the value of later positions — temporal-difference learning avant la lettre.

1960·Ronald Howardcontrol

Dynamic Programming and Markov Processes. Policy iteration: alternate evaluation and greedy improvement. Modern actor-critic methods are its sampled, incremental descendants.

1960·Rudolf Kalmancontrol

A New Approach to Linear Filtering and Prediction Problems, with the linear-quadratic regulator alongside — the exactly solvable case of Bellman’s problem and the benchmark learning controllers are measured against. Pontryagin’s maximum principle (English translation 1962) supplies the continuous-time counterpart, developed independently in the USSR; Kalman & Bucy (1961) add the continuous-time filter, and Stratonovich and Zakai carry filtering nonlinear.

1960·Holt, Modigliani, Muth & Simoncontrol

Planning Production, Inventories, and Work Force. The Carnegie project: quadratic costs give linear decision rules for production and inventory — LQ control before the name, estimated on factory data. The incubator of the next entry.

1961·John Muthequilibrium

Rational Expectations and the Theory of Price Movements. Expectations “are essentially the same as the predictions of the relevant economic theory.” Verified in the text: Muth cites HMMS chapters 2–4 and 19, and his centerpiece example is a commodity market with inventory speculation. Rational expectations was born inside inventory control.

1962·Edmund Phelpsportfolio

The Accumulation of Risky Capital. Ramsey’s control problem made stochastic and sequential by dynamic programming. The verified, directly cited parent of both 1969 lifetime-portfolio papers.

1965·Samuelson & McKeanportfolio

Rational Theory of Warrant Pricing (Industrial Management Review), with McKean’s appendix A Free Boundary Problem for the Heat Equation. American-option pricing posed and solved as optimal stopping, eight years before Black–Scholes; Samuelson & Merton (1969) put it inside a utility equilibrium — the bridge to 1973.

1965·David Blackwellcontrol

Discounted Dynamic Programming. The discounted MDP of Bellman and Howard put on rigorous measure-theoretic footing, with Blackwell optimality — the theorem base under every convergence guarantee later claimed in RL.

Control enters finance, 1965–1990

1965·Karl Åströmcontrol

Optimal Control of Markov Processes with Incomplete State Information. The POMDP: the belief state is a sufficient statistic, and dynamic programming runs over it. His adaptive-control program is the control-side twin of RL — Sutton, Barto & Williams later call RL “direct adaptive optimal control.”

1966·Miller & Orrcontrol

A Model of the Demand for Money by Firms. Firm cash under a random walk with fixed transfer costs: do nothing inside a band, jump to a return point on hitting either edge — the two-threshold-with-return-point policy that reappears in commodity storage bands and dealer inventory bounds. Eppen & Fama (1969) — Fama doing inventory control — and Constantinides (1976, 1978) make the bands rigorous.

1968·Harold Demsetzdealers

The Cost of Transacting. Immediacy is a costly service and the bid–ask spread is its price — the founding paper of market microstructure economics, in the reference lists of Stoll (1978) and Grossman–Miller (1988).

1969·Paul Samuelson & Robert Mertonportfolio

Lifetime Portfolio Selection by Dynamic Stochastic Programming and Lifetime Portfolio Selection under Uncertainty: The Continuous-Time Case, back to back in REStat 51(3). Samuelson solves lifetime consumption and portfolio choice by discrete-time DP, building on Phelps and Mossin (1968); Merton brings Itô calculus and the Hamilton–Jacobi–Bellman equation into finance, yielding the Merton fraction.

1971·Robert Mertonportfolio

Optimum Consumption and Portfolio Rules in a Continuous-Time Model. The general many-asset theory; the 1973 ICAPM turns the same machinery into equilibrium asset pricing with intertemporal hedging demands.

1971·Lucas & Prescottequilibrium

Investment Under Uncertainty. Rational-expectations industry equilibrium shown to be the solution of a planner’s dynamic program — the template “equilibrium = value function of a control problem,” built explicitly on Muth.

1971·Paul Samuelsonequilibrium

Stochastic Speculative Price. The price of a storable commodity “determined as the solution to a stochastic-dynamic-programming problem”; proves the competitive speculative equilibrium coincides with the planner’s optimum, in Gustafson’s setting.

1971·Bagehot & Smidtdealers

Two Financial Analysts Journal pieces stake out the field’s two mechanisms: The Only Game in Town — “Walter Bagehot,” pseudonym of Jack Treynor — has dealers losing to informed traders and recouping from liquidity traders; Smidt has the dealer as an active price-setter managing inventory. The two lines meet in 1985.

1973·Black, Scholes & Mertonportfolio

The Pricing of Options and Corporate Liabilities and Theory of Rational Option Pricing. Dynamic replication: the delta hedge is a feedback control law that renders the payoff attainable, collapsing pricing to a PDE. Harrison–Kreps (1979) and Harrison–Pliska (1981) then recast it by duality: no-arbitrage means a martingale measure exists.

1973·David Jacobsoncontrol

Optimal Stochastic Linear Systems with Exponential Performance Criteria. Risk-sensitive (LEQG) control: exponential-of-quadratic cost, certainty equivalence fails, and the optimal controller equals a deterministic game. Whittle (1981, 1990) completes the theory; the same exponential criterion returns in Hansen–Sargent robustness and in risk-sensitive asset management.

1976·Mark Garmandealers

Market Microstructure. Coins the term; a dealer facing Poisson order arrivals sets quotes to avoid gambler’s ruin — inventory first enters as a viability constraint. The origin point of the quoting strand; Benston & Hagerman (1974) had supplied the OTC spread empirics.

1976·Magill & Constantinidesportfolio

Portfolio Selection with Transactions Costs. Proportional costs added to Merton’s problem and the no-trade cone conjectured; Constantinides (1986) shows investors optimally tolerate wide drift from target.

1977·Ian Wittenrl

An Adaptive Optimal Controller for Discrete-Time Markov Environments. Contains the earliest known published temporal-difference rule — in a control journal. Sutton and Barto discovered it only in 1981, while finishing the actor-critic work.

1977·Kydland & Prescottequilibrium

Rules Rather than Discretion. Optimal control of policy fails when the controlled agents solve their own forward-looking problems: the policymaker’s problem carries the private sector’s equilibrium response as a constraint, making discretionary plans time-inconsistent.

1978·Hans Stolldealers

The Supply of Dealer Services in Securities Markets. The spread as compensation for a risk-averse dealer’s displacement from his desired portfolio — risk-bearing cost, not market power.

1978·Robert Lucasequilibrium

Asset Prices in an Exchange Economy. A representative agent’s Bellman equation prices all assets. Cox, Ingersoll & Ross (1985) supply the continuous-time counterpart: an HJB equation generating the equilibrium term structure.

1980·Amihud & Mendelsondealers

Dealership Market: Market-Making with Inventory. Builds on Garman: bounded inventory, quotes adjusted dynamically as functions of the inventory state, a preferred inventory position.

1981·Ho & Stolldealers

Optimal Dealer Pricing Under Transactions and Return Uncertainty. The canonical stochastic-DP dealer: price-sensitive Poisson arrivals (adopted from Garman), return diffusion, risk aversion; inventory shifts quote levels, not spread. The paper the entire modern quoting literature descends from — then a 27-year gap in the line.

1982·Bensoussan & Lionscontrol

Impulse Control and Quasi-Variational Inequalities (with the companion variational-inequalities volume). The PDE machinery for (s, S)-type policies: the value function of an impulse problem characterized as the solution of a QVI. Harrison, Sellke & Taylor (1983) solve the clean Brownian case; Korn (1998) carries the machinery into Merton’s problem.

1983·Barto, Sutton & Andersonrl

Neuronlike Adaptive Elements That Can Solve Difficult Learning Control Problems. The actor-critic architecture — a TD critic training a trial-and-error actor — balancing a pole. Develops Klopf’s (1972) hedonistic-neuron program.

1985·Glosten & Milgrom; Kyledealers

Bid, Ask and Transaction Prices in a Specialist Market and Continuous Auctions and Insider Trading. The adverse-selection alternative made formal: Glosten–Milgrom’s spread exists at zero inventory, explicitly naming Ho–Stoll, Amihud–Mendelson and Garman as the contrast; Kyle’s insider solves a dynamic control problem while the market maker filters order flow — control plus filtering inside one equilibrium.

1986·Gennotte; Detempleportfolio

Optimal Portfolio Choice Under Incomplete Information (with Detemple’s equilibrium version the same year). Filtering formally enters portfolio choice — Kalman-filter the drift, then optimize on the estimate — made unavoidable by Merton’s 1980 demonstration that drifts are what markets never let you measure. Lakner (1995, 1998) makes it rigorous.

1988·Grossman & Millerdealers

Liquidity and Market Structure. Demsetz’s immediacy made equilibrium: the number of market makers set by the cost of continuous presence against the value of immediacy to outside investors; among its cited antecedents, Cohen–Maier–Schwartz–Whitcomb’s proof (1981) that transaction costs alone sustain a spread.

1988·Richard Suttonrl

Learning to Predict by the Methods of Temporal Differences. TD(λ) defined and named, with convergence in mean for linear features; credits Samuel, Witten, and the Widrow–Hoff rule as antecedents.

1989·Christopher Watkinsrl

Learning from Delayed Rewards. The thesis that married trial-and-error learning to Bellman–Howard dynamic programming, introducing Q-learning. Watkins & Dayan (1992) prove convergence under exactly the Robbins–Monro step-size conditions. Stokey & Lucas’s Recursive Methods (1989), the same year, codifies DP as the language of economic equilibrium.

1990·Davis & Normanportfolio

Portfolio Selection with Transaction Costs. The Magill–Constantinides problem solved rigorously as singular stochastic control: reflect the portfolio at the boundary of the no-trade wedge. Shreve–Soner (1994) redo it with viscosity solutions.

Learning meets markets, 1992–2001

1992·Williams; Tesaurorl

REINFORCE: unbiased Monte-Carlo policy gradients — optimize the policy directly, not through a value function. And TD-Gammon: TD(λ) plus a neural network plus self-play reaches master-level backgammon, the existence proof cited by both the DQN and AlphaGo papers.

1992·Deaton & Laroqueequilibrium

On the Behaviour of Commodity Prices (and JPE 1996). The Gustafson–Samuelson storage equilibrium estimated on thirteen commodity series, via the numerical revival by Wright & Williams (1982) and Scheinkman & Schechtman (1983). The DP-generated price function explains spikes and skewness; autocorrelation less well.

1994·Jaakkola, Jordan & Singh; Tsitsiklisrl

On the Convergence of Stochastic Iterative Dynamic Programming Algorithms and Asynchronous Stochastic Approximation and Q-learning. Q-learning and TD proved convergent by embedding them in Robbins–Monro stochastic approximation — the 1951 foundation made load-bearing.

1996·Bertsekas & Tsitsiklisrl

Neuro-Dynamic Programming. RL systematized as approximate dynamic programming for the control and OR communities — the idiom through which it entered financial engineering. Tsitsiklis & Van Roy (1997) delimit convergence with function approximation.

1996·Malaney, Ilinski & Youngportfolio

Gauge theory reaches economics and finance: Malaney’s Harvard thesis builds a gauge connection for index numbers (the Malaney–Weinstein connection); Ilinski (1997; book 2001) reads arbitrage as the curvature of a discounting connection; Young (1999) renders FX as a lattice gauge theory. None of it touches dealer markets: as of August 2026 no published work applies gauge or symmetry language to market making — the gap the final entry on this page occupies.

1997·Moody & Wurl

Optimization of Trading Systems and Portfolios, then Moody & Saffell, Learning to Trade via Direct Reinforcement (2001). Recurrent RL learns the trading policy directly, maximizing a differential Sharpe ratio rather than forecasting prices — the founding direct-RL trading papers.

1998·Bertsimas & Loexecution

Optimal Control of Execution Costs. Dynamic programming for buying a block over a fixed horizon under linear impact; the risk-neutral answer is the even split. Opens the execution branch.

1999·Bielecki & Pliskaportfolio

Risk-Sensitive Dynamic Asset Management (with Fleming & Sheu, 2000–2002, in parallel). Jacobson–Whittle’s exponential criterion arrives in portfolio theory as factor-based allocation trading long-run growth against variance; Björk, Davis & Landén (2010) consolidate the partial-information line.

2000·Almgren & Chrissexecution

Optimal Execution of Portfolio Transactions. Risk aversion added to Bertsimas–Lo (cited directly): an efficient frontier of liquidation schedules trading impact cost against volatility risk, with permanent versus temporary impact. Obizhaeva & Wang (working paper 2005) later add order-book resilience.

2000·Routledge, Seppi & Spattequilibrium

Equilibrium Forward Curves for Commodities. The storage DP enters mainstream finance: forward curves from a Deaton–Laroque-style equilibrium, with the inventory non-negativity constraint creating an embedded timing option and endogenous convenience yield.

2001·Chan & Sheltondealers

An Electronic Market-Maker (MIT AI Memo). The first reinforcement-learning market maker: policy-gradient and TD agents quoting in a simulated Garman-style order-flow market, balancing profit and inventory. Cited by both Nevmyvaka–Feng–Kearns and Spooner et al.

2001·Longstaff & Schwartz; Tsitsiklis & Van Royportfolio

Valuing American Options by Simulation and Regression Methods for Pricing Complex American-Style Options. Approximate dynamic programming arrives in finance proper: regress continuation values along simulated paths — approximate value iteration for the Wald-lineage stopping problem. Rogers (2002) supplies the martingale dual. Hansen & Sargent (2001, book 2008) meanwhile import H∞ robust control into economics.

Convergence, 2003–2013

2003·Robin Hansondealers

Logarithmic Market Scoring Rules (working paper 2002; with “Combinatorial Information Market Design,” 2003). An automated market maker as a scoring rule: log-sum-exp cost function, bounded worst-case loss — quoting without a quoting desk. Chen & Pennock (2007) recast it as inventory-utility; Othman et al. (2010) make its liquidity adaptive.

2004·Kakade, Kearns, Mansour & Ortizexecution

Competitive Algorithms for VWAP and Limit Order Trading. Worst-case guarantees from online algorithms — a genuinely separate computer-science tradition (its reference list contains no Bertsimas–Lo, no Almgren–Chriss), merged with the control tradition only by the next entry.

2006·Nevmyvaka, Feng & Kearnsexecution

Reinforcement Learning for Optimized Trade Execution. The first large-scale empirical RL for execution — Q-learning over millisecond order-book states on NASDAQ data — and the verified bridge: it cites Bertsimas–Lo, Almgren–Chriss, the VWAP line, and Chan–Shelton. The lead in Halperin’s Quora answer that seeded this page; Nevmyvaka went on to head machine-learning research at Morgan Stanley.

2006·Huang, Malhamé & Caines; Lasry & Lionsequilibrium

Nash Certainty Equivalence (2006) and Mean Field Games (2007), independently. A continuum of agents each solving a control problem against the population distribution; equilibrium is a coupled backward-HJB / forward-Kolmogorov system — equilibrium-as-control made literal. Carmona & Delarue (2018) give the probabilistic codification.

2008·Avellaneda & Stoikovdealers

High-Frequency Trading in a Limit Order Book. Ho–Stoll 1981 revived for electronic books — the paper says so itself (“closely related to a paper by Ho and Stoll”) — with exponential-utility reservation prices and fill intensities λ(δ) = Ae−kδ. Ends the 27-year gap in the quoting line.

2013·Guéant, Lehalle & Fernandez-Tapiadealers

Dealing with the Inventory Risk. Inventory bounds turn the Avellaneda–Stoikov HJB into linear ODEs: the first verification theorem and closed-form quotes as functions of inventory. Guilbaud & Pham (2013) add market orders; Cartea, Jaimungal & Penalva’s 2015 textbook makes it the standard curriculum.

2013·Gârleanu & Pedersenportfolio

Dynamic Trading with Predictable Returns and Transaction Costs. The synthesis in closed form: Markowitz’s statics, Merton’s dynamics, and transaction costs resolved with the LQ regulator — aim in front of the target.

2013·Mnih et al.rl

Playing Atari with Deep Reinforcement Learning (Nature version 2015; AlphaGo 2016). Watkins’ Q-learning stabilized with deep networks, experience replay and target networks — citing Watkins–Dayan and Tesauro. The enabling technology for everything below; DDPG (2015) and PPO (2017) supply the continuous-action line finance mostly uses.

Into production, 2014–2026

2014·Peter Cottonjpm

Theoretically motivated width and skew for OTC trading, and the discovery of a key gauge symmetry making control and RL feasible — in production at J.P. Morgan, begun on arrival in 2013 and live the following year. The theory circulated as Cotton & Papanicolaou, Trading Illiquid Goods (working paper; presented at NYU, 2018); the arXiv record of the mid-2010s work is the final entry on this page. With the assistance of Nick West, Erik Clossen and Pascal Tomecek.

2017·LOXMjpm

J.P. Morgan’s deep-RL execution engine, reported July 2017 as live in European equities since Q1, trained on billions of historic and simulated transactions.

2017·Man AHLindustry

A learning order router in live production, per Risk.net’s fund-of-the-year citation: trades allocated across internal algos, dealer algos and the human desk by Thompson sampling — a 1933 bandit algorithm running money. Bloomberg had reported Man Group’s machine-learning models live across $12.7bn of funds that July.

2017·Ritter; Halperinrl

Machine Learning for Trading: Q-learning with a mean–variance reward discovers statistical arbitrage under costs — the paper that made RL respectable in the practitioner quant literature. And QLBS: option price and hedge emerge as the optimal Q-function of a risk-adjusted MDP — Black–Scholes done by Q-learning. Kolm & Ritter’s 2019 survey maps the field onto the RL formalism.

2018·Buehler, Gonon, Teichmann & Woodjpm

Deep Hedging (Quantitative Finance 2019). A derivatives book hedged by deep RL under transaction costs and risk limits, sidestepping complete-market pricing — written inside J.P. Morgan’s equities QR. In production hedging vanilla index flow from 2018; by 2022, per Risk.net, 70% of the bank’s Euro Stoxx flow options were quoted and hedged by machine, with Deep Bellman Hedging (2022) the actor-critic sequel. Buehler left for the market maker XTX in 2022.

2018·Bacoyannis, Glukhov, Jin, Kochems & Songjpm

Idiosyncrasies and Challenges of Data Driven Learning in Electronic Trading. The bank’s own candid account of RL in production trading: thousands of micro-decisions, reward-shaping difficulties, regulatory constraints — “we backtest, therefore we overfit.” The same year, Manuela Veloso founds J.P. Morgan AI Research, whose ABIDES simulator (2019) and multi-agent RL dealer markets (2019) carry the simulation line forward.

2018·Spooner, Fearnley, Savani & Koukorinisdealers

Market Making via Reinforcement Learning. TD agents with an asymmetrically dampened inventory reward recover skewing behavior on high-frequency order-book data — citing both lineages: Ho–Stoll and Avellaneda–Stoikov on one side, Chan–Shelton and Nevmyvaka–Feng–Kearns on the other. Cardaliaguet & Lehalle’s mean field game of controls lands MFGs in microstructure the same year.

2019·Guéant & Manziukdealers

Deep Reinforcement Learning for Market Making in Corporate Bonds. Actor-critic networks replace finite-difference HJB solvers for multi-asset OTC quoting — the point where the Ho–Stoll control lineage and the RL lineage formally merge, on exactly the asset class of the 2014 entry.

2019·Guo, Hu, Xu & Zhangequilibrium

Learning Mean-Field Games. MFG equilibrium computed by model-free Q-learning, with guarantees; fictitious play (Perrin et al. 2020) and scalable deep RL (Laurière et al. 2022) follow. With Gomes & Saúde’s price-formation MFG (2021) — the price as a Lagrange multiplier on flow balance for a stored commodity — the loop closes: Gustafson’s 1958 storage equilibrium, solved by reinforcement learning.

2020·RBC Aidenindustry

Aiden, built with Borealis AI — whose RL program RBC launched in 2017 with Richard Sutton as head academic advisor — goes live: a client-facing deep-RL VWAP execution agent, the clearest non-JPM production deep-RL claim on record, followed by Aiden Arrival (2022) and a deep-hedging stress-testing pilot (2021). Beyond the banks, public production claims stay rare: the proprietary market makers publish almost nothing, and XTX frames its billion-euro ML build-out as supervised forecasting.

2020·Tooling and textbooksrl

Dixon, Halperin & Bilokon’s Machine Learning in Finance (Springer), the first graduate text with a full part on RL; and FinRL, the open-source library that became the field’s dominant experimental ecosystem.

2020·Constant-function market makersdealers

Uniswap’s x·y = k deployed at scale (v2 whitepaper, 2020; formal analysis Angeris et al., 2019–21): market making reduced to a fixed convex curve — no inventory control at all, the degenerate opposite of Ho–Stoll. Milionis, Moallemi, Roughgarden & Zhang (2022) price the consequence — loss-versus-rebalancing, the CFMM’s Glosten–Milgrom adverse-selection bill — and ZeroSwap (2023) starts handing the curve learned, inventory-style control back.

2023·Hambly, Xu & Yangrl

Recent Advances in Reinforcement Learning in Finance (Mathematical Finance 33(3)). The standard mathematical survey, connecting RL back to the stochastic-control formulations of execution, market making and portfolio choice; the Annual Review of Statistics survey (2025) closes the period, naming explainability and robustness as the barriers to adoption.

2026·Peter Cottondealers

On a Simple Relationship Between Order Imbalance, Skew and Width in Over-The-Counter Trading. The imbalance–skew–width symmetry isolated from the mid-2010s market-making work — this project. The dealer strand of this page, begun with Garman’s viability constraint, restated as a symmetry.

1 Harris published the square-root formula in Factory, The Magazine of Management in 1913. In 1931 F. E. Raymond’s book — the first on the subject — garbled the citation, making the original untraceable; R. H. Wilson’s 1934 Harvard Business Review article then popularized the result, which operations research thereafter called the “Wilson formula.” The original was rediscovered only in 1988, and Erlenkotter’s “Ford Whitman Harris and the Economic Order Quantity Model” (Operations Research 38(6), 1990) restored priority.

Citation links were verified against primary sources where linked; the companion literature map covers the inventory-and-pricing control strand in depth.