Seventy years of portfolio construction compressed into three regimes: the policy portfolio, the robo-advisor, and now the AI agent. This piece maps the evolution, surveys what today's AI offerings actually do — versus what their marketing says — and frames the question the rest of this series will attack: how do you prove a model is behaving inside your risk tolerance?
The story of modern wealth management is usually told as continuous progress. It reads better as three distinct regimes, each defined by what the industry chose to automate — and each ending when its central assumption broke.
Harry Markowitz's 1952 "Portfolio Selection" gave the industry its operating system: diversification as a computable trade-off between expected return and variance. ERISA (1974) institutionalized it, and by the 1980s the 60/40 stock/bond mix had become the default policy portfolio for pensions and, through balanced funds, for households. The track record justified the default. A US 60/40 blend delivered roughly 8.3% annualized over the trailing 30 years with single-digit volatility, and — measured at calendar year-ends since 1928 — never posted a negative 10-year return. The human advisor sat on top of this machinery, charging roughly 1% of assets largely for allocation, discipline, and hand-holding.
The regime's central assumption was the stock-bond hedge. 2022 broke it: with inflation forcing rates up, a globally diversified 60/40 fell about 16% — by some measures the strategy's worst year since 1937, and the first modern year in which both stocks and bonds fell by double digits simultaneously. The portfolio wasn't dead (it returned roughly +30% cumulatively over the following two years, and Vanguard still defends the construct while tilting its 2026 model toward bonds), but the episode ended 60/40's claim to being a complete answer. It also created the market opening every AI pitch now walks through: if the static allocation can fail, perhaps the allocation should think.
Betterment (founded 2008, launched 2010) and Wealthfront (launched as a robo in 2011) automated the policy portfolio rather than rethinking it: risk questionnaire in, MPT-optimized ETF portfolio out, with threshold rebalancing and — from 2012 — automated tax-loss harvesting. The innovation was economic, not intellectual. At ~25 basis points versus ~100 for a human advisor, robo-advice quadrupled the addressable market and forced fee deflation across the industry.
Two things are worth being precise about, because they shape how to read today's AI claims. First, robo-advisors were never AI — they were deterministic rules engines executing 1950s portfolio theory. No learning, no language, no inference. Second, the pure-play economics mostly failed. Goldman's Marcus Invest was sold to Betterment (2024); JPMorgan, UBS, and US Bank all shut their robos (2024–2025); BlackRock wound down FutureAdvisor (2023). Wealthfront survived to a December 2025 IPO — but its S-1 revealed that roughly three-quarters of revenue now comes from cash-management spread, not advisory fees. The survivors are those attached to a giant (Vanguard Digital Advisor at ~$312B including hybrid programs, Schwab, Empower) or those that found a different monetization. Automation compressed fees faster than it created margin.
ChatGPT's November 2022 launch — the same month the 60/40 was completing its worst modern year — marks the boundary. What changed was not portfolio math but the interface and the scope: models that read documents, hold conversations, personalize at scale, and increasingly act. The adoption curve since has been steep on the sell side (98% of Morgan Stanley advisor teams on the firm's GPT-4 assistant; 95% of wealth and asset managers with generative AI scaled into multiple use cases per EY's 2025 survey) and surprisingly steep on the retail side (62% of surveyed US retail investors have used AI to inform decisions). The frontier moved again in 2025–2026 with agentic execution: Robinhood's sandboxed agentic-trading accounts (100,000+ funded since May 2026) and broker APIs like Alpaca's MCP server, where LLM agents — not humans — now drive order flow.
Each regime automated the previous one's scarce resource. The policy portfolio automated judgment into math; the robo automated the math into product; AI is now automating the conversation, the research, and — at the frontier — the decision itself. Which is precisely why the control question, not the capability question, is where this series will spend most of its time.
This series will ultimately treat institutional allocators, advisory practitioners, and retail investors as three separate discussions — the requirements barely overlap. As an opening frame, here is what each constituency is actually buying today, what it must demand, and where its specific risk sits. Later pieces expand each column into its own evaluation.
Research synthesis at scale (AlphaSense, Hebbia, Rogo), portfolio-analytics copilots (Aladdin Copilot, MSCI AI Portfolio Insights, Bloomberg Document Insights), and genuine ML alpha generation inside quant and multi-strategy funds (Bridgewater's ~$2B AIA fund, Point72's Turion).
Model inventory and validation lineage; explainability sufficient for an investment committee; diligence on foundation-model dependence (a vendor's model-version bump is a model change); data licensing and audit trails.
Backtest credibility and crowding. A 2026 survey of LLM-trading studies found only 2 of 19 had clean time-consistent data splits; regulators' emerging worry is correlated agent behavior — herding at machine speed.
The productivity layer: meeting capture and CRM automation (Jump, Zocks, Morgan Stanley Debrief, Merrill's Meeting Journey), plan generation (Conquest), proposals (Powder), next-best-action (TIFIN AG), and household-level agents (Savvy Intelligence).
FINRA 3110 supervision extended to AI output; books-and-records for prompts and drafts; client consent for recording; human review before anything reaches a client; Marketing-Rule discipline on what you claim your AI does.
Paying for AI without a data foundation — 64% of wealth firms lack a unified data layer, and ROI remains "elusive" per F2 Strategy's 2026 survey. And fiduciary duty cannot be delegated to a model: the advisor owns every recommendation the machine drafts.
Portfolio digests and screening (Robinhood Cortex), holistic advisory recommendations (PortfolioPilot), AI-built custom indexes (Public's Generated Assets), and — newest — sandboxed agents that execute trades (Robinhood Agentic Trading).
Know the autonomy level you're granting; know whether "assets on platform" means managed or merely linked; know who executes and who is accountable; verify AI claims — the SEC's first AI-washing cases hit consumer-facing advisers.
Confidently wrong output and unaudited performance marketing. Studies put LLM error rates on personal-finance questions near 35%; "70% win rate" claims from AI stock-picker vendors are unaudited; and the "self-directed" legal framing shifts responsibility for agent behavior onto the user.
The most useful single axis for evaluating any "AI investing" product is not model quality — it is how much decision authority the system holds. Nearly everything shipped at scale today sits on the bottom two rungs. The interesting engineering, and all of the interesting risk, lives in the climb.
Deterministic MPT allocation, threshold rebalancing, tax-loss harvesting. The robo-advisor stack. No inference.
Digests, summaries, research Q&A: Schwab Portfolio Insights, Bloomberg Document Insights, Aladdin Copilot, meeting notetakers. Human decides everything.
Personalized, actionable recommendations the user executes: PortfolioPilot, Arta AI, Cortex trade ideas, Conquest's plan engine. This is where fiduciary and suitability questions become live.
Agents execute within hard deterministic constraints: Robinhood's sandboxed pre-funded agentic accounts, LLM agents on Alpaca's MCP rails, advisor-supervised rebalancing (Vise). The frontier of what ships to retail in 2026.
Model holds the mandate. Exists today only inside quant/multi-strat funds with institutional risk infrastructure (Bridgewater AIA, Point72 Turion, Voleon). Not currently permissible as a retail advisory product — "RIAs aren't allowed to hire an AI agent to manage money… it doesn't suit the rules currently" (TradePMR's Robb Baldwin).
Below is a working map of the offerings that define the space as of August 2026 — filterable by market segment and by rung on the autonomy ladder. Two honest caveats baked into the data: platform-reported "assets" for retail AI tools generally mean linked or analyzed assets rather than discretionary AUM, and several vendors listed have unaudited performance claims or enforcement history, flagged inline.
| Offering | Segment | Autonomy | What it actually does | Scale |
|---|
Scale figures are the most recent company-reported or press-reported numbers as of Aug 2026; bases differ (AUM vs linked assets vs users vs adoption) and are labeled per row. Hover a row for sourcing and caveats.
Three patterns stand out. First, the shipped product is overwhelmingly L1–L2. The scaled deployments — Morgan Stanley's assistant, JPMorgan's LLM Suite, Merrill's Meeting Journey, Schwab's Portfolio Insights — are drafting and explanation layers with a human firmly in the loop. Discretion remains confined to quant funds and sandboxes. Second, the guardrail pattern is converging across every serious deployment: consent-gated inputs, retrieval restricted to curated corpora, human review before client contact, sandboxed and pre-funded execution, "self-directed" legal framing. Firms discovered the same containment architecture independently because the regulatory physics is the same. Third, the economics rhyme with the robo era. Budgets are surging while measured ROI stays thin (F2 Strategy, 2026), notetaker startups are raising dueling Series Bs while platforms absorb their feature set, and at least one pure-play AI robo (Q.ai) has already shut down. The capability is real; the business models are still sorting.
"Does it generate alpha?" is the question everyone asks. "How do I know it's behaving inside my risk tolerance — and how would I prove it?" is the question that determines whether any of this is investable. Two facts frame the answer in 2026, and both are uncomfortable.
Fact one: the regulatory scaffolding just got thinner, not thicker. The SEC withdrew its only AI-specific rule proposal (the predictive-data-analytics conflicts rule) in June 2025. In April 2026, the Fed, OCC, and FDIC replaced SR 11-7 — the model-risk bible since 2011 — with shorter, principles-based guidance that explicitly excludes generative and agentic AI from scope, promising a future request for information instead. FINRA's position is technology-neutral ("the rule is the rule, no matter how the method changes"), the EU AI Act does not classify investment advice as high-risk, and its high-risk deadlines are themselves slipping under the Digital Omnibus. Meanwhile every SEC AI enforcement action to date — Delphia, Global Predictions, Rimar — punishes lying about AI, not AI misbehaving. The burden of making AI safe inside a mandate has been left, almost entirely, to the firms deploying it.
Fact two: the raw models are not trustworthy enough to skip the engineering. On FinanceBench — real questions against real SEC filings — GPT-4-Turbo paired with a retrieval system failed 81% of the time; with idealized retrieval, accuracy jumped to ~89%. That ~70-point swing is the single most important number in AI wealth management: it says the control point is the retrieval and validation architecture around the model, not the model choice. The same lesson generalizes to execution: a 2026 survey of LLM-trading research found essentially none of the published performance claims auditable (2 of 19 studies with clean data splits, 1 of 19 with transaction costs).
The architecture that answers the risk-tolerance question is old, proven, and borrowed from electronic trading: the stochastic model proposes; a deterministic layer disposes. Knight Capital's $460M, 45-minute self-destruction in 2012 — and the Market Access Rule (15c3-5) enforcement that followed — established the template: hard pre-trade checks, independent kill paths, deployment controls, all sitting outside the thing being controlled. Applied to an AI manager, the stack looks like this:
Notice what this stack implies for the buyer's diligence question. "Is your AI good?" is unanswerable. "Show me the deterministic constraint set my mandate compiles into, your eval suite and its failure rates, your last champion-challenger promotion decision, and the kill path that doesn't depend on the agent's own cooperation" — that is answerable, and today almost no retail-facing product will answer it. Closing that gap is where this series goes next. The Bank of England is already asking the systemic version of the same question: with half of finance firms running agentic AI, Deputy Governor Sarah Breeden noted in June that "our frameworks were not built to contemplate autonomous agents" — and that human-in-the-loop for every agent action is "unlikely to be realistic." The controls have to be structural.
Part I set the map. The proposed continuation — each piece standalone, each building the evaluation framework the finale needs:
Hands-on evaluation of PortfolioPilot, Cortex, Public's Generated Assets, Magnifi and peers: identical prompts and test portfolios, scored on recommendation quality, consistency across sessions, risk-profile adherence, and what happens when you push against the guardrails.
From notetaker to next-best-action: where the ROI actually shows up, the data-layer prerequisite, supervision and books-and-records obligations, and a build-vs-buy framework for RIAs — expanding the practitioner lens from §02.
The MCP-broker rails (Robinhood, Alpaca), sandbox design, and the core engineering question: how an IPS becomes a machine-readable constraint set — encoding risk tolerance as tracking-error budgets, VaR caps, concentration limits, and drawdown triggers the agent cannot argue with.
Proving behavior, not asserting it: eval design and hallucination benchmarks, deflated-Sharpe and CPCV backtest hygiene, champion–challenger promotion, shadow trading, drift monitoring, and model-version change management — the answer to "how do I know it's true."
The post-SR-11-7 vacuum, SEC/FINRA posture, EU AI Act drift, state regimes, and the BoE kill-switch debate — maintained as a living reference page, updated as the 2026–2027 rulemaking cycle resolves.
Proposed functionality enhancements: what a platform designed around the control stack — rather than retrofitted with it — should look like, from continuous suitability to portfolio-level agent orchestration, and which incumbents are closest.
Companion deliverables per piece: an interactive dashboard for kineticalpha.com and a LinkedIn distribution post, consistent with prior Kinetic Alpha research releases.