Skip to content
TRADE FLOW RESEARCH

Trade Flow Opportunity Engine

Choosing what to import is normally intuition with a spreadsheet behind it. This engine reads public trade statistics, ranks the product categories a country under-imports relative to comparable economies, then discounts every candidate by the things that actually kill an import business: local production, tariff walls and regulatory burden.

  • DomainInternational trade and sourcing
  • EngagementInternal research engine
  • StageRunning engine, operated in-house
  • Economic Complexity
  • Public Trade Data
  • Analytical Engine
  • Decision Support

Public statistics, one store, one ranked brief

Trade statistics are published openly by several international bodies, and almost nobody uses them to answer a buying question, because each source speaks a different format, reports to a different latest month and counts in different units. The engine unifies them into one analytical store, computes published economic-complexity measures on top of them, and ranks under-served import categories into a brief a person can act on.

It is honest about what it is. The analytical core works and a full-scale run has been completed end to end. That run produced a ranked brief and a machine readable companion file. It is not a hosted product: there is no deployment, no multi-user story, and no validated commercial outcome behind the ranking. It has a documented method, and a defect it found in its own output and corrected structurally.

THE CHALLENGE

Sourcing decisions made on hearsay

The information needed to choose an import category exists and is public. The reason it goes unused is that nothing turns it into a decision.

The category is chosen before it is checked

A category gets picked because someone heard it was moving. The supplier is found first and the disqualifiers arrive afterwards, once money and time are already committed.

The things that kill a deal surface last

Local production, tariff walls and regulatory burden decide whether an import lane is viable at all, and they are the last facts anyone checks rather than the first filter anyone applies.

Open data that nobody reconciles

The statistics are public, but each source publishes in its own format, to its own latest month, in its own units. Reconciling them by hand for every candidate category is not work a person does twice.

THE SOLUTION

An engine that ranks, and explains the ranking

One store behind many sources, published methods computed on top of it, a score built from multiplicative gates, and a brief that shows its own arithmetic.

Multiple public sources behind one interface

International trade statistics, macro development indicators, economic-complexity measures and a gravity dataset for distance and trade-agreement effects, all landing in one analytical store that every model reads from.

  • Each source isolated, so one being down or blocked degrades that source alone
  • An ingestion log recording which sources landed and which degraded on every run
  • Cached fetching, so re-running an analysis does not re-pull the world

Published economic methods, implemented and tested

The measures the economic-complexity literature uses, implemented as published, computed against the unified store, and exercised offline in tests.

  • Revealed comparative advantage, its symmetric form, and revealed import advantage
  • Product-space density and the fitness-complexity iteration
  • A gravity model comparing expected trade against observed trade

A score built from gates, not weights

Strength signals average into one core score. The disqualifiers do not join that average, they multiply it, so a category the country already produces or deliberately walls out cannot be floated by strength elsewhere.

  • Peer comparison against a basket of comparable economies, each on its own trailing window
  • Domestic production and tariff posture applied as multiplicative gates
  • Every point of the final score attributable to a named signal, exposed beside it

A brief that shows its own calculation

The output is a ranked, sector by sector brief with a machine readable companion file, and it closes by explaining exactly how each number was produced.

  • Regulatory context carried beside the ranking, because a gap in a regulated category is a different opportunity
  • A closing section giving the formulas, the peer basket and the materiality floor
  • An explicit statement that the ranking is an application-layer heuristic, with no published index behind it
BEHAVIOUR

What it does when the data is not there

Most of the interesting decisions in this build are about what happens when the evidence is missing.

No stage is allowed to raise

Every stage returns a status. A missing dependency, an unreachable service, a failed subprocess or an empty result becomes a named outcome with a hint for fixing it.

Insufficient data is an answer

The drivers were built before the data existed and default to doing no network at all. They report insufficient data, and never a number that nothing supports.

Missing evidence is unclear, never zero

A source that is blocked records unclear. Writing zero would let an absence of evidence read as evidence of absence, which is how a ranking quietly lies.

Where the data cannot be trusted, it asks a person

One input had no source worth trusting, so the run stops and asks the operator. The answer is recorded as manually sourced.

THE IMPACT

What the build actually produced

No commercial outcome is claimed here, because none was measured. What follows is what changed in the engine and in the quality of what it writes.

Multiplicative gates

Ranking correctness

The first version put protected categories at the top, precisely the ones a country walls out with its own tariffs, because the disqualifiers were weighted terms in a sum. Turning them into multiplicative gates made that class of error structurally impossible.

Shows its work

Auditable output

Every brief ends with the formulas, the peer basket and the materiality floor, so a reader can disagree with a rank on the method.

Degrades, not fails

Run reliability

A blocked or unavailable source is logged and the run continues on what landed. The pipeline reports insufficient data, and never presents a partial result as a complete one.

Last reviewed:

Have a decision buried in public data?

If the numbers that should drive a decision are already public and nobody has the time to reconcile them, that is the kind of engine we build. Tell us what the decision is.

Start a conversation