Data

The datasets behind The Matcha Index articles are published here so anyone can verify our numbers or build on them. Files are updated when we re-collect; each filename carries its collection date.

Region and label claims read from product pages (1,123 listings, 16 sellers)

This supersedes the two files below for any question about what a seller discloses. The earlier collections read claims from sellers’ machine-readable product feeds. That was wrong: Shopify’s feed carries a body_html field, and for several sellers the specification block a shopper actually reads — Origin, Cultivar, Production Year, Net weight, Shelf life — is rendered into the page and never enters the feed. Measured on one seller’s matcha, 2 of 14 listings named a growing region in the feed and 14 of 14 named one on the page. Every disclosure figure in the earlier files is understated, always in the direction that made sellers look less forthcoming.

This file reads the rendered page instead, with navigation and footer removed by frequency: any six-word run appearing on more than 60% of one seller’s own pages is boilerplate, except a labelled specification such as Origin: or Harvested:, which repeats because the seller states it on every product. All 1,116 pages read are cached, so any figure can be re-derived from the exact text it came from.

Download CSV — japanese_green_tea_pages_2026-08-27.csv

1,518 excluded rows with the reason for each — teaware, packaging, gift cards, and three fruit and flower powders sold under the matcha name that contain no tea.

Download CSV — japanese_green_tea_pages_excluded_2026-08-27.csv

The collection and classification script, with 29 documented guards. Each guard records the wrong number that prompted it, including the two that produced this file.

Download Python — collect_japanese_green_tea_pages_2026-08-27.py

Matcha brand entry and top tins (265 listings, 12 US shops)

The tin-level extract behind the matcha brands comparison — every matcha listing from the August 26, 2026 survey after removing Yunomi (yen prices), sweetened and flavoured products, sets and warehouse duplicates: 265 rows with price, weight, per-gram price, stock on the survey date, a size-window flag (20–40 g, bulk, out of band) and a culinary-label flag. Two tins discussed in the article are absent here because the sellers’ feeds carried no weight for them (Arteao’s On The Go, Rishi’s Everyday Matcha); how each was handled is in the article’s method note.

Download CSV — matcha_brands_survey_2026-08-26.csv

Region names on matcha labels (1,123 listings, 16 sellers)

1,123 tea listings across 2,386 size variants from the same 16 sellers, re-collected August 27, 2026 with an expanded place-name vocabulary. This is the dataset behind the region-name article. It supersedes the August 26 file for place-name questions only: the earlier run’s vocabulary held 24 prefecture names and no municipalities, which meant Nishio and Wazuka were counted from product titles alone while every other place was counted from titles and descriptions. The vocabulary now holds 38 names including the Uji-area towns, and the counts for those two places roughly tripled.

Download CSV — japanese_green_tea_2026-08-27.csv (superseded for disclosure questions — read from the feed, not the page)

1,518 excluded rows, with the reason for each — teaware, packaging, gift cards, and three fruit and flower powders sold under the matcha name that are not tea.

Download CSV — japanese_green_tea_excluded_2026-08-27.csv

The collection and classification script, with 26 documented guards. Each guard records the wrong number that prompted it.

Download Python — collect_japanese_green_tea_2026-08-27.py

Japanese green tea survey (1,125 listings, 16 sellers)

1,125 tea listings across 2,391 size variants from 16 sellers — 15 American or American-facing brands plus Yunomi, a Japanese marketplace of small farms — collected August 26, 2026. One row per size variant: seller, currency, title, the seller’s own product_type, tea type assigned, price, net weight, price per gram, and the disclosure flags used in the article (region, cultivar, shading, harvest, caffeine claims, flavour vocabulary, brewing figures).

Download CSV — japanese_green_tea_2026-08-26.csv (superseded for disclosure questions — read from the feed, not the page)

1,514 excluded rows, with the reason for each. Teaware, packaging, gift cards and other non-tea items are excluded on the seller’s own product_type field before any figure is computed. Counting exclusions is impossible from the main file alone, so they are published separately and can be audited rather than taken on trust.

Download CSV — japanese_green_tea_excluded_2026-08-26.csv

The collection and classification script, including the 24 documented guards — compound-word handling for bancha, multipack weight parsing, the product_type-over-title rule, and the pagination fix that stopped the largest seller’s feed being truncated.

Download Python — collect_japanese_green_tea.py

Matcha price dataset

292 products and size variants from the official online stores of 12 brands (13 storefronts), collected August 24, 2026: brand, product, price, net weight, price per gram, labeled grade, and named cultivar where disclosed.

Download CSV — matcha_prices_2026-08-24.csv

Buyer-review dataset

1,762 verified-buyer reviews published on the official stores of Rocky’s Matcha, Naoki Matcha, and Nekohama (February 2020 – August 2026), collected via the stores’ public review widgets: brand, product, rating, review title, review text, date. Capped at the most recent 100 reviews per product by the widget API.

Download CSV — matcha_reviews_okendo_2026-08-24.csv

Sencha disclosure audit (299 listings, 7 sellers)

Every listing matching sencha or 煎茶 in the public product feeds of fifteen shops, collected August 26, 2026 — seven of them sell it. That is 299 distinct listings and 717 size variants, coded for the steaming grade named (asamushi, chuumushi, fukamushi), the cultivar, the harvest, the growing region, process claims, caffeine claims, flavour vocabulary, brewing instructions, and price per gram. Fifty-three matches were excluded, most of them ceramics — “sencha cup” is a standard vessel size in Japan. Used in Sencha: What It Is, What It Costs, and What 299 Listings Say About How It Was Steamed.

Download CSV — sencha_disclosure_2026-08-26.csv
Download the collection script — sencha_disclosure.py

Hojicha disclosure audit (143 listings, 15 sellers)

Every listing matching hojicha, houjicha or hōjicha in the public product feeds of fifteen sellers, collected August 26, 2026: 143 distinct listings and 330 size variants, coded for the base tea named, harvest timing, roast level and method, roast date, caffeine claims, growing region, flavour vocabulary, brewing instructions, and price per gram. This measures what sellers publish, not what is in the bag. Used in Hojicha: What It Is, What It Costs, and What 143 Listings Say It’s Made Of.

Download CSV — hojicha_disclosure_2026-08-26.csv
Download the collection script — hojicha_disclosure.py

Two things in that dataset were sorted by hand rather than by pattern, because the wording does not separate them. Both files carry the quote each call rests on, so the calls can be disagreed with:

Download CSV — hojicha_bancha_calls_2026-08-26.csv — the ten listings mentioning bancha, split into the six claiming their own tea is bancha and the four describing what hojicha usually is.
Download CSV — hojicha_harvest_calls_2026-08-26.csv — the seventy listings using a harvest or season word, split into the fifty-nine stating when their own leaf was picked and the eleven using the word for something else.

Label disclosure audit (133 listings, 11 brands)

Every matcha listing from eleven direct-to-consumer brands, coded for whether the product description names a harvest period, cultivar, growing region, shade-growing method, milling method, or any date. Collected August 25, 2026. This measures disclosure, not accuracy — we did not verify any brand’s claims. Used in How to Tell If Matcha Is High Quality.

matcha_label_disclosure_2026-08-25.csv — 133 rows

matcha_label_terms_2026-08-25.json — the exact term lists used to code each trait, so you can see what did and didn’t count

collect_label_claims.py — the collection script, including every exclusion rule

analyze_review_vocab.py — the review vocabulary analysis

Matcha set dataset (57 sets, 9 shops)

Every matcha preparation set carried by twelve direct-to-consumer Japanese tea sellers, collected August 25, 2026, together with the individual components those same shops sell separately. Used in What a Matcha Set Actually Costs.

matcha_sets_curated_2026-08-25.csv — 73 candidate sets with a verdict column. The twelve we excluded are kept in the file with the reason written out, so you can disagree with any of them.

matcha_set_contents_2026-08-25.csv — what each seller states is in the box, parsed from its own contents list

matcha_components_2026-08-25.csv — 1,749 standalone components (whisks, bowls, scoops, sifters, stands, tins, cloths). One caveat about this file: its cat column is a coarse automatic label applied from product titles, and it misfiles a minority of rows — a kyusu sold with a strainer lands under “sifter,” and some futaoki lid rests land under “scoop.” Every component price quoted in the article was matched to a named product by hand, not taken from a category average. Treat the column as a rough index, not a classification.

mass_market_sets_2026-08-25.csv — the 29 set listings Google Shopping displayed for “matcha set” on August 25, 2026, recorded by hand. This is a snapshot of a personalised, rotating panel rather than a sample we controlled, and it is the weakest source on this page. It is published so the figures in the article can be checked, not because it is representative.

collect_matcha_sets.py — the collection script

curate_matcha_sets.py — the inclusion and exclusion rules, with every exclusion reason in the source

parse_set_contents.py — how each seller’s contents list was read

Genmaicha and its base teas (1,440 listings, 8 sellers)

Every genmaicha, matcha-iri genmaicha, sencha, bancha, kukicha, hojicha and gyokuro listing carried by eight sellers of Japanese tea, collected August 25, 2026. Built to answer one question: genmaicha is roughly half rice, so is it cheaper than the tea it is blended from? Used in What Genmaicha Actually Costs.

genmaicha_prices_2026-08-25.csv — 1,440 price-and-size combinations. Columns include kind (which tea), form (leaf / powder / teabag), grams, and cur, since one seller prices in yen. 1,381 rows publish a weight; the other 59 cannot be reduced to a per-gram figure and appear in no calculation.

collect_genmaicha_prices.py — the collection script. Its header sets out the known failure modes in this kind of listing data and the guard applied to each, which is also the fastest way to see what the classifier can still get wrong.

One caveat we found the hard way: an earlier run of this script counted fifteen “Matcha-Infused Genmaicha” listings as plain genmaicha, and counted a deep-steamed sencha as genmaicha because it shared a product page with one. Both are fixed in the file above, but they are the reason the kind column should be treated as a machine’s best reading of a product title rather than a seller’s own categorisation.

Gyokuro: shading disclosure and flavour vocabulary (65 products, 6 sellers)

Gyokuro is the one Japanese tea defined by a number — about twenty days under shade, below which it is classed as kabuse tea instead. These two files record whether sellers state that number, and what words they use in place of it. Used in What Gyokuro Actually Costs. Prices for the same products are in the tea-type file above.

gyokuro_shade_2026-08-25.csv — every gyokuro product at six sellers, with the shading period where stated and the exact sentence it was read from (evidence), so you can check our parsing rather than trust it. The kind column separates karigane and kabuse, which carry the word “gyokuro” without being it.

gyokuro_flavour_2026-08-25.csv — which flavour words appear in each of those 65 descriptions. This counts marketing copy, not flavour: nobody has checked whether the teas described as umami taste more savoury than the others.

collect_gyokuro_shade.py — the collection script, including the exclusion rules.

Two errors this file has already produced, both found by reading titles rather than by running a check: a stem tea named “Gyokuro Kukicha” survived an exclusion list that only looked for “karigane,” and a teapot survived one that looked for “teapot” while the product was called a “Tea Pot.” Both sat in the “doesn’t mention shading” bucket, which is where a non-tea item naturally lands. Removing them moved the product count from 67 to 65. The versions above are corrected; the episode is why the evidence column exists.

Rishi: buyer reviews and three years of price history

Built for Is Rishi Matcha Still Good?, which tests a claim circulating on Reddit that Rishi’s matcha declined around 2024.

reviews_structured_2026-08-26.csv — 12,281 buyer reviews from Rishi, Rocky’s Matcha and Naoki Matcha, 2012 to August 2026: shop, product, rating, date, whether the reviewer was a verified buyer, and whether the review covers a product or the store as a whole. The share of verified buyers is the column that matters most: when a shop switches from reviews people volunteer to reviews it requests from every customer, ratings fall without the product changing.

Review text is not republished here — it belongs to the people who wrote it, and reposting ten thousand strangers’ words in bulk is a different act from linking to them. The consequence is that the breakdown of what the twenty low reviews complain about cannot be reproduced from this file alone. The collection script below will pull the text directly if you want to check that table.

rishi_price_history_2026-08-26.csv — what four Rishi matchas cost at eight archived dates between September 2023 and July 2026, read from Internet Archive snapshots of Rishi’s own product pages. Failed snapshot fetches are recorded as missing rather than as zero, which is why some rows have an empty price.

okendo_reviews.py — the review collection script. Its header sets out the known failure modes in this kind of review data, including the one that matters most here: rows a shop files as store-wide feedback about delivery are not reviews of a tea, and mixing them in moves the numbers.

rishi_price_history.py — the price-history script. Reading an old price off an archived page is harder than it looks, because a product page also carries the prices of whatever else the shop was cross-selling that day; the header explains how the script proves a price belongs to the product being measured.

Jade Leaf: catalogue, three shelves and seven years of price history

Three files behind the Jade Leaf piece, collected on August 27, 2026. The brand matters to this index because it is the one an American shopper is most likely to meet in a supermarket, and because a third of what it sells under the word matcha is a drink mix rather than tea.

jadeleaf_listings_2026-08-27.csv — 48 variant rows covering all 31 products the shop lists, with the shelves the shop itself files each one on, its specification rows, its cultivar row, and the ingredient and nutrition fields it publishes in structured data. Three columns are worth reading before the rest. on_pure_shelf and on_mix_shelf come from the shop’s own collection pages and never from words in a product name, because “Matcha Latte Mix” and “Ceremonial Matcha” both contain the word. on_nav_matcha_shelf records the shop’s general Matcha shelf, which carries products from both sides and so cannot be used to classify anything. And net_grams is the size the seller states in the variant name, not Shopify’s grams field, which is a shipping weight. Two free-text columns present in our working copy — the product description and the seller’s Q&A block — are not in the published file, because the Q&A carries buyers’ own questions and we would rather quote from it than republish it in bulk.

jadeleaf_shelves_2026-08-27.csv — 67 rows: the same products priced on the brand’s own site, on the Walmart marketplace storefront the brand operates itself, and at Target, each row carrying the listing title as shown, the size as stated, the identifier we matched on and the page we read it from. Retailer prices move and a marketplace listing is not a price guarantee, so read this as what those three sites displayed to us in one session. Amazon is absent: the browser used resolved it to a non-US storefront and returned prices in the wrong currency. One row pair is worth looking at directly — the two 1.06 oz Teahouse tins on the Walmart storefront carry different barcodes and different prices, and nothing in their titles distinguishes them.

jadeleaf_price_history_2026-08-27.csv — 239 priced rows from 78 Internet Archive captures of the four pure-matcha product pages, the oldest from July 2019. No capture failed to load and no page was missing its price block, so this is the whole of what the archive holds for these four products. Prices are recorded per variant, never averaged across sizes, because this shop sells the same tea in up to five sizes at once and the sizes did not move together: between mid-2025 and today its retail packs rose between 0.2% and 9.7% while all three of its 1 lb bags rose between 11.1% and 20.8%. Captures that sit below their neighbours are left in rather than filtered out — we cannot tell a promotion from a repricing with certainty, and the file lets you decide.

jadeleaf_listing.py — the catalogue collector. Its header lists the failure modes this kind of listing data has and the guard applied to each. Two are specific to this shop and worth knowing about: the specification rows live in a theme block and the ingredient figures in a JSON-LD block, so neither appears in the product feed and every disclosure field is read from the page instead; and the shop writes its cultivar label three different ways (Tea Cultivars, Cultivars, Tea Cultivar), so reading only the plural form finds six of the eight rows and drops the products that name exactly one variety.

jadeleaf_history.py — the archive collector. It reads prices only from the structured product block belonging to the page being measured, identifies that page by its handle where the theme provides one and by its canonical link where it does not, and records a failed capture as missing rather than as zero.

Kettl: 45 matcha, 12 growers and seven years of prices

Two files behind the Kettl piece, collected on August 28, 2026. This shop is the opposite end of the market from Jade Leaf, and the interesting thing about it is not its prices but what it prints next to them: a grower, a village, a cultivar row and a harvest year on every tin.

kettl_listings_2026-08-28.csv — 291 variant rows covering all 183 products the shop lists, of which 45 are matcha. Read is_matcha before anything else. It is 1 only where the shop’s own matcha shelf and its own product type agree, because this shop’s two classification systems disagree with each other: its matcha-green-tea shelf also carries two hojicha powders, a subscription and a tea class, and its wholesale shelf is named wholesale-matcha-houjicha-powder-bulk and holds both teas at once. Taking the shelf alone gives 49; requiring both gives 45. The four disagreements are flagged in shelf_type_disagree rather than deleted, so you can see them. is_wholesale separates the 1 kg bags, which would otherwise drag every per-gram figure down. net_grams is the size the seller states, in the variant name or in its own Packaging row; the shopify_grams_field column is kept beside it only to show that the two disagree, because Shopify’s figure is a shipping weight.

kettl_price_history_2026-08-28.csv — 71 rows from 54 Internet Archive captures of seven product pages, the oldest from July 2019. A capture that shows two sizes produces two rows, which is why the row count is larger than the capture count. Eight teas were queried; the eight were chosen because they had archived pages at all, so teas that have been on the shelf longest are over-represented and this is a selection rather than a sample. One of the eight, Tsuji Family Blend, returns nothing even under the shop’s real handle and is absent. Three others first returned nothing because we had guessed their URLs from their product names — a zero here means the archive has nothing only after the shop’s own handle has been tried. Both this file and the listing file contain some Japanese: those are the shop’s own product names (福寿抽茶, 爽香, さみどり抽茶), not stray notes.

kettl_listing.py — the catalogue collector. Its header lists the failure modes this kind of listing data has and the guard applied to each. The one specific to this shop is guard 3: what counts as matcha comes from two of the seller’s own fields having to agree, never from the word “matcha” in a product name — this shop sells matcha bowls, matcha whisks, matcha chocolate and matcha classes.

kettl_history.py — the archive collector. It reads prices only from the structured product block belonging to the page being measured, identifies that page by its handle where the theme provides one and by its canonical link where it does not, and records a failed capture as missing rather than as zero.

Use and citation

Free to use for any purpose with attribution: “The Matcha Index (thematchaindex.com)”, with a link. If you spot an error, email info@thematchaindex.com and we’ll correct the dataset and note the change here.

Method in one paragraph

Prices and weights come from each brand’s public store feed, collected on the date in the filename, covering every listed product (no sampling). Weights listed in ounces are converted at 1 oz = 28.35 g. Listings without a computable net weight are included in the file but excluded from per-gram statistics. Reviews come from the brands’ public review widgets and are verified-buyer entries as labeled by those systems.

Hojicha price dataset (280 listings, 14 sellers)

Every hojicha listing we could reach from thirteen US brands and Yunomi, the Japanese marketplace, normalised to dollars per gram. Powder and loose leaf are separated. Yunomi prices are in yen; the cur column marks which. Collected August 25, 2026. Used in What Hojicha Powder Actually Costs.

hojicha_prices_2026-08-25.csv — 291 rows

collect_hojicha_prices.py — the collection script, including the currency check and every exclusion rule

Bamboo whisk (chasen) price dataset (72 listings, 11 sellers)

Every bamboo matcha whisk we could reach from eleven sellers, with the specifications each listing states — prong count, bamboo type, Takayama/Nara origin, handmade claim, shin construction. Collected August 25, 2026. Yunomi prices are in yen; the cur column marks which. Used in What a Matcha Whisk Costs.

chasen_prices_2026-08-25.csv — 72 rows

collect_chasen_prices.py — the collection script, including every exclusion rule and the word-boundary matching that keeps whisky and Cat’s Whiskers tea out of the set

Naoki: listings, name history and reseller prices

Everything Naoki Matcha publishes about each tin it sells, what those tins were called and cost in the past, and what the same tin was listed for elsewhere under a different name. Collected August 26, 2026. Used in One Tin, Two Names.

naoki_listings_2026-08-26.csv — 33 rows: every product and size variant, with the seller’s own specification rows (origin, grade, harvest, tasting note, serving size), its own shop categories, price per gram, and a count of where the words ceremonial, superior, premium and culinary appear — product name, spec row, description, category and image alt kept in separate columns, because the same word means different things in each. Two columns are worth reading before the rest: shopify_grams_field is the shop’s shipping weight and reports 45 g for a 40 g tin, which is why net_grams is taken from the size printed on the tin instead; and cultivars_named_on_page is an exact-string test against the eighteen cultivar names listed in the script, so it will miss any spelling not in that list.

naoki_listing_history_2026-08-26.csv — 105 rows from 44 Internet Archive captures of eight product pages between April 19, 2024 and May 6, 2026, giving 104 price readings and no failed fetches. Coverage is uneven and the file shows it: Superior Blend and Organic First Spring are captured twelve times each, Ujitawara seven, Chiran once, and Barista Pro Blend not at all.

naoki_resellers_2026-08-26.csv — 5 rows. The listings Google Shopping returned for the exact phrase naoki matcha superior ceremonial blend on one day, with each seller’s posted price beside Naoki’s own. A Shopping panel is a rotating sample of offers rather than a census, and the largest seller under that name — Naoki’s own Amazon listing — is absent because its US price was not readable from where we searched. Read the file as what one shopper was shown on August 26, 2026.

naoki_listing.py — the listing collector. Its header sets out the failure modes this kind of data has, the first of which is the one that matters most here: a shop’s product feed is not its page. This shop keeps its specification table in a theme block the feed never returns, so counting words in the feed alone reports a term as absent when the page shows it. Every field here is read from the rendered page.

naoki_history.py — the archive collector, which reads price and product name only from the structured block whose own handle matches the page being measured, and records a failed snapshot as missing rather than as zero.