01 / THE PROPER ROLE OF AI
Separate research acceleration from decision authority.
An AI system can summarize filings, compare language across thousands of documents, classify events and help analysts formulate alternative explanations. These are meaningful research advantages, but they are not the same as knowing an asset’s future return. Markets contain changing incentives, incomplete data and reflexive behavior. A confident output can still be based on stale evidence, a weak proxy or a relationship that no longer holds.
The most useful division of labor is therefore asymmetrical. Machines handle scale, consistency and retrieval; people remain responsible for framing the question, judging source quality, identifying missing context and deciding whether a conclusion is fit for use. Human review should not be a ceremonial approval at the end. It should challenge the model at the points where assumptions enter the process.
Use AI to increase the number and quality of questions a research team can test—not to make uncertainty disappear.
NIST’s AI Risk Management Framework describes AI risk management as a lifecycle activity organized around governance, context mapping, measurement and management. Applied to investment research, that means a model should be evaluated not only for statistical performance but also for where it sits in a decision, who can override it and what happens when its evidence is unavailable or contradictory.
02 / A CONTROLLED WORKFLOW
Move from source evidence to an auditable conclusion.
A credible AI-assisted workflow preserves the path from raw information to the final research judgment. The goal is not to automate every step. It is to make each transformation visible enough to inspect, reproduce and challenge.
Define the question
State the decision, time horizon, investable universe and conditions that would disprove the thesis before selecting a model.
Build the evidence set
Prefer traceable primary sources, preserve publication timestamps and record revisions so the system cannot learn from information that was unavailable at the decision time.
Generate competing views
Ask the system to surface supporting evidence, contradictory evidence and plausible alternative explanations instead of optimizing for a single persuasive narrative.
Connect to economics
Translate a signal into a possible cash-flow, discount-rate, liquidity or positioning channel. A label without a transmission mechanism is not yet an investment thesis.
Test implementation
Evaluate turnover, transaction costs, market impact, liquidity, capacity and portfolio interaction—not only a model’s isolated prediction score.
Record the decision
Document the evidence used, the human owner, the override logic, the risks accepted and the monitoring conditions that could change the conclusion.
This sequence also helps distinguish two very different failures. A model failure occurs when the system behaves outside its tested limits. A use failure occurs when a sound tool is applied to the wrong question, market or decision. Both require controls.
03 / VALIDATION DISCIPLINE
Test the full decision chain, not only the model.
Validation begins with the data lineage: what entered the system, when it became available and how it was transformed. It then asks whether the model is stable across periods, market environments and reasonable alternative specifications. Finally, it tests how errors propagate into portfolio decisions. A small classification error can become material if it repeatedly pushes a concentrated position in the same direction.
| Validation layer | Core question | Useful evidence |
|---|---|---|
| Data | Was every input knowable at the time? | Source log, timestamp policy, revision history, missing-data treatment. |
| Model | Does performance survive realistic variation? | Out-of-sample tests, perturbation tests, benchmark comparisons, error analysis. |
| Decision | Does the output improve a defined choice? | Cost-aware simulation, exposure attribution, override records, rejected signals. |
| Operations | Can the process fail safely? | Access controls, monitoring thresholds, fallback procedure, incident log. |
Independent challenge matters. The team that builds a model has deep context but may also be invested in its success. Validation should have enough authority, technical skill and separation to question design choices, reproduce results and restrict use when evidence is weak. The Federal Reserve’s model-risk guidance is written for supervised institutions, but its distinction among sound development, effective challenge and governance offers a useful process reference for any model-dependent research organization.
A model can be statistically accurate and still be unsuitable for a specific portfolio, horizon or liquidity condition. Fitness for use is a decision-level judgment.
04 / GOVERNANCE AND CONTROLS
Govern the claim, the model and the action.
Good governance begins before deployment. Each model or AI-enabled workflow should have a named owner, a documented purpose, clear prohibited uses and a review frequency proportionate to its impact. Material changes to data, prompts, model versions or portfolio use should trigger renewed testing.
Minimum control record
- Purpose and boundary: what the system is designed to do—and what it must not be used to decide.
- Source provenance: where evidence comes from, how it is timestamped and how corrections are handled.
- Performance limits: known error modes, unstable environments and the benchmark that defines useful improvement.
- Human authority: who approves use, who can override an output and who can suspend the system.
- Monitoring: drift, data outages, unusual output patterns, concentration effects and realized decision quality.
- Vendor dependency: version changes, service interruption, data retention, access rights and exit plans for third-party tools.
Public language is part of governance too. SEC officials have warned against misleading claims about the use of artificial intelligence. A responsible description should match the actual workflow and controls. Terms such as “AI-driven” should not imply autonomous accuracy, regulatory endorsement or guaranteed performance.
05 / INVESTOR QUESTIONS
Five questions cut through most AI claims.
- What decision does the system improve? A clear answer should name the task, horizon and comparison baseline.
- What evidence can be traced? Outputs should connect to sources, timestamps and transformations that can be reviewed.
- How was the claim tested? Look for out-of-sample evidence, realistic costs and disclosure of failed or rejected approaches.
- Who remains accountable? There should be a human owner with authority to challenge, override or stop the workflow.
- What happens when conditions change? A credible process defines monitoring thresholds, fallback procedures and revalidation triggers.
These questions do not guarantee a good investment outcome. They do make it easier to distinguish a controlled research capability from a marketing label.
06 / SOURCES AND FURTHER READING
Primary references behind this guide.
This article is an original synthesis. The following public materials inform its governance, validation and disclosure framework; they are linked so readers can examine the underlying guidance directly.
- 01NIST AI Risk Management Framework 1.0↗
A voluntary framework for governing, mapping, measuring and managing AI risks across the lifecycle.
- 02Federal Reserve SR 26-2: Model Risk Management↗
Supervisory guidance covering model development, validation, governance, third-party dependencies and risk-based application.
- 03SEC Statement on AI Washing↗
A public reminder that statements about an organization’s use of AI should be accurate and not materially misleading.