Guide

AI Quant: Data and Validation

When people put AI into A-share quant, the bottleneck is almost never the model — it is two things: can the data be fetched reliably, on demand, by the AI; and can the conclusion be validated (instead of fooled by a backtest). This page covers both paths and the four traps.

In one line

AI does not solve stock picking; it solves consistency between data and validation. A model can write code, compute indicators, screen in bulk and summarise filings — but where the data comes from, whether the conventions are right, and whether the conclusion is overfitted decide the outcome. Exposing your data as tools the AI can call (MCP), then validating by statistical thresholds, are the two highest-ROI steps.

Two common misreadings

Common misreadings
  • "A stronger model picks better stocks" — the model cannot see data you do not have; data quality sets the ceiling
  • "A high backtest return means a good strategy" — overfitting, survivorship bias and ignored costs can all make a fake look great
  • "The AI will just find the data online" — scraped pages are unstable and irreproducible, so results cannot be re-verified
A more realistic approach
  • Turn data into tools first: let the AI Agent call structured quotes/financials/money flow, instead of reading web pages
  • Fix the conventions before running numbers: fields, adjustment method, trading-day alignment, cost assumptions
  • Decide by threshold: enough samples, confidence-interval lower bound above 0, out-of-sample direction consistent — all three

Where the data comes from: four paths

Four common paths; cost and stability differ a lot (ordered by "can it support long-running work", not by popularity):

Open-source library (e.g. AkShare)
Free · Medium
Research, prototypes, academic use (docs state "for academic research")
Docs: endpoints "often need maintenance because target sites change" — keep upgrading the package; official risk note: mind commercial risk. An official HTTP API variant (AKTools) also exists
Points-based service (e.g. Tushare)
Points threshold (tiered, not consumed) · High
Structured historical data, tidy schemas
Regular data needs 5000+ points for higher frequency; minute data requires a separate permission; you still wrap it for your AI Agent
Multi-source API (this service)
Free tier · paid from ¥9.9 · High (multi-source failover + circuit breaker)
Long-running systems, AI Agent and review pipelines
No minute bars, no full-text news/announcements (bring your own)
Scraping pages yourself
Highest time cost · Low
One-off needs with no alternative source
Breaks on any redesign; anti-bot and legal risk; not reproducible

How to choose: for exploration, a free open-source library is fine; once it must run long-term or be called by an AI Agent automatically, switch to an interface built for reliability (failover, explicit rate limits, stable fields) — otherwise your time keeps leaking into data repair.

Connecting the data to an AI Agent: two paths

Same data, two integration styles with different division of labour:

Path 1: REST (you write the fetching code)

You control timing and parameters, then hand the result to the model. Good for backtests, batch computation, scheduled jobs. Free endpoints need no key; one GET returns data in a unified envelope `{ ok, endpoint, tier, elapsed_ms, source, data }`.

import requests

r = requests.get(
    "https://api.ashareapi.com/v1/kline",
    params={"code": "sh600667", "period": "day", "count": 60},
    timeout=10,
)
rows = r.json()["data"]          # newest first; rows[0] is the latest bar
# When feeding a model, pass only the fields it needs -- not the whole payload
Path 2: MCP (configure once, the AI Agent fetches)

MCP (Model Context Protocol) is the standard way AI Agents call external tools. Configure it once and, when you ask about quotes in chat, the Agent calls the tool itself — no fetching code on your side. This is the lowest-friction shape for "let the AI do research".

# Claude Code, for example (add once, works in every later session)
claude mcp add --transport http ashareapi https://api.ashareapi.com/mcp

# Verify: ask "What is the price of 600667 right now?"
# Judge by whether it actually called a tool (e.g. ashare_quote), not by whether the answer looks plausible

Free tools (quote / K-line / hot list / market overview / up-down distribution) need no key; for paid tools add one `Authorization` header. Full config for 19 clients is on the [MCP page](/en/mcp); China platforms are covered in the [China platform guide](/en/docs/mcp-china).

After you have the data: the four validation traps

A backtest can produce whatever conclusion you want. These four are the usual false-positive factories:

Overfitting: tuning until it fits

Parameters that look great in-sample often die out-of-sample. Always hold out OOS: define the rule on the first 70%, adjudicate on the last 30%; if the direction flips, discard it.

Survivorship bias: only survivors left

Backtesting with today's listing is assuming you knew who would not delist. Use the full universe as it was (including delisted/suspended names).

Ignoring costs: net is the real number

Commissions, stamp duty and slippage eat most short-horizon edge. A-share cost magnitudes: 10cm about -0.50%, 20cm about -1.10% (round trip). Compute net first.

Too few samples: three for three means nothing

A small sample of wins has no statistical meaning. Our bar: n ≥ 30 and a confidence-interval lower bound above 0; if the interval contains 0, keep accumulating.

One more: once you fix the data conventions, do not change them. Switching fields, adjustment method or cost assumptions is a different experiment — earlier conclusions no longer apply and the validation has to restart.

FAQ

Can AI predict price moves?

Not reliably. Models are good at bulk processing, writing code, computing indicators, reading filings and turning messy data into structured conclusions. Automating those steps saves a lot of time and emotional trading — but "predicting price" is not within its capability boundary.

Do I need to write code to use an AI Agent for quant research?

With MCP you need no fetching code (one line of config, then just talk). For backtesting and validation you should at least read Python — judging whether a conclusion holds requires understanding the statistics yourself, and that step should not be outsourced to the model.

Is the free tier enough?

For research and light use, yes: five data endpoints are free without a key; anonymous limits are 5/min, up to 60/min after solving one PoW challenge. For bulk backtests or shared egress IPs, add a key (tiered limits; Standard allows 120/min).

Can I backtest directly with it?

Yes for quotes, K-lines, financials, money flow and top-trader boards — but we do not provide minute bars or full-text news/announcements. Keep Tushare or a specialist vendor for those; the two can be used together.

What about fully automated AI trading?

This page is about data and research tooling only, and does not suggest unattended automated trading. Any action involving real money must keep risk control and human confirmation in the loop.

Last updated: 2026-09-21

Endpoints, fields and limits come from live responses of this service and from the MCP `tools/list` (verified 2026-09-21: envelope fields ok/endpoint/tier/elapsed_ms/source/data, 5 free endpoints, anonymous 5/min, 60/min after PoW, Standard 120/min — all checked against the server code); AkShare characteristics are quoted from its official docs (akshare.akfamily.xyz/introduction: the risk note and the "endpoints often need maintenance" statement); Tushare point rules are quoted from its official docs (tushare.pro/document/1?doc_id=108 permissions, doc_id=234 minute data requires a separate permission); the MCP definition is quoted from modelcontextprotocol.io ("open-source standard"); cost magnitudes come from our own backtest convention settings.

Next steps

Connect the data to your AI Agent (one command), or check the endpoint list to see coverage.