Surveying the Vietnamese POS market
A survey of five POS vendors in Vietnam: public sources crawled with Playwright, the data re-checked with Pandas, then charted for comparison.
- Python
- Playwright
- Pandas
- Jupyter
Before advising a shop on which point-of-sale software to buy, I wanted numbers rather than impressions. This project surveys five widely used POS vendors in Vietnam — KiotViet, Sapo, MISA CukCuk, iPOS.vn and PosApp — from public sources, captured on 23 July 2026.
Public data is not clean data#
Vendor pricing pages are written to sell, not to be compared. Sapo lists its Start Up plan at 249,000đ per month, but the 170,000đ figure only applies with a two-year commitment. iPOS.vn publishes FABiBox at a promotional price valid until 30 September 2026, while full FABi carries no public price at all. PosApp's hardware page still carries a promotion note that expired on 30 September 2025.
So most of the work was not collecting numbers. It was recording the conditions attached to each number, so the figure can be audited later.
Crawling that leaves a trail#
crawl_data.py drives Playwright with Chromium, reads its list from
data/sources.csv (17 official sources), checks robots.txt per domain, and opens only
two pages at a time by default. Every URL leaves four artefacts: machine-readable JSON
(metadata, headings, tables, links, JSON-LD), Markdown for me to re-read, a raw HTML or
PDF snapshot, and a prompt packet.
def save_record(record, output_dir, prompt, prompt_max_chars, save_html):
source_id = record["source"]["source_id"]
raw_html = record.pop("raw_html", None)
# Checksum of the extracted text: a later crawl shows exactly which page moved.
record["content_sha256"] = hashlib.sha256(
record.get("text", "").encode("utf-8")
).hexdigest()
json_path.write_text(json.dumps(record, ensure_ascii=False, indent=2), encoding="utf-8")
markdown_path.write_text(render_markdown(record), encoding="utf-8")
prompt_path.write_text(
render_prompt_packet(record, prompt, prompt_max_chars), encoding="utf-8"
)The crawler also flags suspicious snapshots on its own: text shorter than 300 characters,
text that hit the --max-chars ceiling and may be truncated, or a page showing CAPTCHA
markers. Those flags go straight into crawl_manifest.csv and into the prompt packet, so
I never read a broken capture as if it were the real page.
The LLM reads; it does not decide#
The prompts/ directory exists because extraction from the page is handed to an LLM:
Vietnamese POS pricing pages are messy HTML tables with conditions scattered around them,
and rule-based parsing breaks the moment a layout changes. The prompt is tightly
constrained — never infer facts absent from the source, attach a short evidence quote and
the source URL to every fact, keep the original prices, separate list price from
promotional price from monthly-equivalent price, use null where the source is silent,
and tag every field with a confidence level.
Validate before plotting#
app.py runs validate_data before any chart is drawn, and raises rather than plotting
something wrong. The conditions: exactly five vendors with unique keys, every analysis
table covering all five, entry and hardware prices all positive, the capability matrix
matching the published rubric, every source on HTTPS, and price-confidence labels limited
to the two documented values.
def validate_data(tables: dict[str, pd.DataFrame]) -> list[str]:
errors: list[str] = []
# The coverage rubric only allows the five levels written in docs/methodology.md.
feature_values = tables["features"][list(FEATURE_LABELS)].to_numpy(dtype=float)
if not np.isin(feature_values, [0.0, 0.5, 0.75, 0.8, 1.0]).all():
errors.append("Feature coverage values fall outside the documented rubric")
sources = tables["sources"]
invalid_urls = sources.loc[
~sources["url"].astype(str).str.startswith("https://"), "url"
].tolist()
if invalid_urls:
errors.append(f"All source URLs must use HTTPS: {invalid_urls}")
if set(tables["vendors"]["price_confidence"]) - {"Cao", "Trung bình"}:
errors.append("Unexpected price confidence label")
if errors:
raise ValueError("Data validation failed:\n- " + "\n- ".join(errors))
return checksTwo notebooks close the loop: 01_validate_data.ipynb re-runs every condition, and
02_compare_prices.ipynb pins the median and mean with assert. Edit one price cell and
forget to update the report, and the notebook goes red.
Charts and normalisation#
app.py produces three charts: entry software pricing, the reference starter hardware
bundle, and a capability-coverage matrix across eight functional groups. The matrix
deliberately does not collapse into a total score, because the weighting depends on
each shop's business model; it measures how much a vendor publishes, not product quality.
Every chart carries its conditions as a footnote under the figure.
Outcome#
- As of 23 July 2026, public entry software pricing ran from 129,000đ to 270,000đ per month, with a median of 199,000đ and a mean of 197,600đ
- The reference starter hardware bundle fell between 7.00 and 9.14 million đồng, median 7.99 million — but the configurations are not equivalent: the CukCuk Starter combo already includes a year of software
- Two price points were downgraded to "medium" confidence because the source still showed an expired promotion or published too little, and that sits in the data rather than being quietly smoothed over
- Every figure is only valid at the capture date; the report is for building a shortlist and a rough budget, not a substitute for a vendor quote