telecomkz_scraper/tools/capture_screen.py
Iliyas Kyrykbayev 5cb347a44b TelecomKz Analytics Mapper: audit fixes, real event keys, 32 mapped screens
Rebuilt the capture and mapping pipeline after an audit found the simulator's
data could not be trusted:

* Hotspot coordinates never matched the screenshots. Capture now scrolls the
  page over CDP and pastes each frame at the measured scrollY, so image pixels
  and DOM coordinates share one grid by construction.
* Metrics were synthesised (1200 + n*410) and presented as analytics. Numbers
  are now attached only when the catalog has a matching row; metrics.json
  carries a `source` label and the UI says "no data" instead of showing zeros.
* Event interception hooked a connector bridge that never fires. The app posts
  to api.amplitude.com using the legacy form-urlencoded v1 API; the hook now
  reads event_type off the wire. 36 keys are verified as `observed`.
* All device access moved into tools/telecom_cdp.py: dynamic WebView socket
  discovery (the PID was hardcoded), id-matched CDP, measured native geometry.
* Editor edits can now be saved to disk; API failures no longer report success
  from a stale result file; screenId is no longer interpolated into a shell.

Screens went from 7 (with fabricated markup) to 32, all verified: image height
equals map height, no out-of-bounds hotspots, no dead links.

The id_card screenshot has been manually redacted - it showed a national ID.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
2026-08-24 18:16:46 +05:00

725 lines
28 KiB
Python
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

"""
TelecomKz screen capture + hotspot mapping.
Produces, for one screen id:
public/assets/screens/<id>_long.png the stitched screenshot
public/assets/screens/<id>_result.json the capture result (also merged into the app map)
How the stitch works
--------------------
The old implementation swiped blindly and glued frames together with OpenCV
template matching, so the height of the result was unpredictable - while the
hotspot coordinates were computed from the DOM document height. The two never
agreed, which is why hotspots landed off-image.
This version drives the scroll from the page itself over CDP and reads the real
scroll offset back after every step. A frame captured at scrollY lands at row
round(scrollY * scale) of the content band, by construction. So an element at
document offset y always sits at
webview_top + round(y * scale)
in both the image and the hotspot rects. No matching, no drift, no guessing.
"""
import argparse
import io
import json
import sys
import time
from PIL import Image
sys.path.insert(0, str(__import__("pathlib").Path(__file__).resolve().parent))
import telecom_cdp as T
# Measure the page in one round trip; every number the stitcher needs comes from here.
PAGE_GEOMETRY_JS = """
(() => ({
url: location.href,
route: location.pathname,
title: document.title || '',
innerWidth: window.innerWidth,
innerHeight: window.innerHeight,
scrollHeight: Math.max(
document.documentElement.scrollHeight,
document.body ? document.body.scrollHeight : 0
),
scrollY: window.scrollY,
maxScrollY: Math.max(
0,
Math.max(
document.documentElement.scrollHeight,
document.body ? document.body.scrollHeight : 0
) - window.innerHeight
)
}))()
"""
# A cheap fingerprint of "what is rendered right now". Some sections (the ID card,
# the tariff catalogue) paint a skeleton first and fill it a second or two later, and
# capturing then yields a screenshot with no controls on it at all.
PAGE_FINGERPRINT_JS = """
(() => {
const body = document.body;
return JSON.stringify({
n: document.querySelectorAll('*').length,
h: document.documentElement.scrollHeight,
t: (body ? (body.innerText || '').length : 0),
skeleton: document.querySelectorAll(
'[class*="skeleton"], [class*="loader"], [class*="loading"], [class*="spinner"]'
).length
});
})()
"""
DOM_ELEMENTS_TEMPLATE = """
((strictOcclusion) => {
const selectors = [
'button', 'a[href]', '[role="button"]', '[onclick]',
'.menu-list-item', '.extra-menu__card', '.bonuses-card',
'.user-balance-card', '.customer-account-card__item',
'[class*="banner"]', '[class*="promo"]', '.swiper-slide',
// Side-drawer rows (Профиль телеком / Помощь / Оферта). They are plain DIVs with
// no role or href, so nothing else here matches them and the whole drawer was
// invisible to the mapper until this selector was added.
'.nav-link'
];
const seen = new Set();
const out = [];
document.querySelectorAll(selectors.join(', ')).forEach(el => {
const style = window.getComputedStyle(el);
if (style.display === 'none' || style.visibility === 'hidden' || style.opacity === '0') return;
const r = el.getBoundingClientRect();
if (r.width < 20 || r.height < 15) return;
const text = (el.innerText || el.getAttribute('aria-label') || el.getAttribute('title') || '')
.trim().replace(/\\s+/g, ' ');
if (!text) return;
// Skip anything an overlay is covering. With the side drawer open the dashboard
// behind it is still laid out and would otherwise be mapped as clickable, so the
// element under its own centre point has to actually be this element.
// Occlusion test. What counts as "covered" depends on how the screen is captured:
//
// strict (single-viewport capture): probe the centre of whatever part is on
// screen, so a card half below the fold is still tested. Needed for the side
// drawer, which covers content that straddles the fold.
//
// lenient (scrolling capture): only test elements fully inside the viewport.
// Elements below the fold get scrolled into view later and are captured then;
// judging them against a sticky banner at scroll 0 would wrongly discard them.
const vx0 = Math.max(r.left, 0), vx1 = Math.min(r.right, window.innerWidth);
const vy0 = Math.max(r.top, 0), vy1 = Math.min(r.bottom, window.innerHeight);
const fullyVisible = r.top >= 0 && r.bottom <= window.innerHeight &&
r.left >= 0 && r.right <= window.innerWidth;
if (vx1 > vx0 && vy1 > vy0 && (strictOcclusion || fullyVisible)) {
const hit = document.elementFromPoint((vx0 + vx1) / 2, (vy0 + vy1) / 2);
if (hit && !el.contains(hit) && !hit.contains(el)) return;
}
// A card and its inner link report the same box; keep the outermost only.
const key = [Math.round(r.left), Math.round(r.top + window.scrollY),
Math.round(r.width), Math.round(r.height)].join(':');
if (seen.has(key)) return;
seen.add(key);
out.push({
text: text.slice(0, 80),
tag: el.tagName.toLowerCase(),
className: String(el.className || '').slice(0, 120),
domId: el.id || '',
dataEvent: el.getAttribute('data-event') || el.getAttribute('data-analytics') || '',
cssLeft: r.left,
cssTop: r.top + window.scrollY,
cssWidth: r.width,
cssHeight: r.height
});
});
return { elements: out };
})
"""
def dom_elements_js(strict_occlusion):
return "(" + DOM_ELEMENTS_TEMPLATE + ")(" + ("true" if strict_occlusion else "false") + ")"
# Back-compat for callers that import the constant directly (live_auto_recorder).
DOM_ELEMENTS_JS = dom_elements_js(False)
# Text -> Amplitude event key. These are the mappings the analyst confirmed; anything
# else gets a provisional key flagged with `keyConfidence: "guessed"` so it is obvious
# in the editor which rows still need a real key.
# Ordered most-specific first: card labels concatenate their title and subtitle, so
# "Подключить интернет Скидки и бонусы" must hit the internet rule before the bonus one.
EVENT_RULES = [
(("подключить интернет",), "CONNECT_INTERNET_CLICK", "Подключение интернета", None),
(("turbo",), "TURBO_CLICK", "Подключение Turbo-скорости", None),
(("telecom shop",), "SHOP_CLICK", "Telecom Shop", None),
(("tv+",), "TV_PLUS_CLICK", "Переход в «TV+»", None),
(("aitu music", "музык"), "MUSIC_CLICK", "Переход в «Музыка»", "music_screen"),
(("лицевой счет", "лицевой счёт"), "ACCOUNT_SELECTOR_CLICK", "Выбор Лицевого Счета", None),
(("свободные средств",), "BALANCE_CARD_PAY_CLICK", "Быстрая оплата баланса", "payments_screen"),
(("мои бонусы", "бонус"), "BONUSES_CLICK", "Раздел «Мои бонусы»", None),
(("детализац",), "DETAILS_CLICK", "Переход в «Детализацию»", "details_screen"),
(("мои услуги", "услуг"), "SERVICES_CLICK", "Переход в «Мои услуги»", "services_screen"),
(("трафик",), "TRAFFIC_CLICK", "Просмотр «Трафик»", "traffic_screen"),
(("платеж", "оплат", "пополн"), "PAYMENTS_CLICK", "Переход в «Платежи»", "payments_screen"),
(("заявк",), "ORDERS_CLICK", "Раздел «Заявки»", "orders_screen"),
(("удв. лич", "удостовер"), "IDENTITY_CLICK", "Удостоверение личности", None),
(("qr",), "QR_PAY_CLICK", "Оплата по QR", None),
(("сервис",), "SERVICES_CATALOG_CLICK", "Каталог сервисов", None),
(("баланс",), "BALANCE_CARD_PAY_CLICK", "Быстрая оплата баланса", "payments_screen"),
]
def slugify_event_key(text):
"""
Provisional key for an unmapped element. Transliterated, because Amplitude event
keys are ASCII - the old code emitted Cyrillic keys like CLICK_ЧАТЫ that could
never match a real event.
"""
table = {
"а": "A", "б": "B", "в": "V", "г": "G", "д": "D", "е": "E", "ё": "E", "ж": "ZH",
"з": "Z", "и": "I", "й": "Y", "к": "K", "л": "L", "м": "M", "н": "N", "о": "O",
"п": "P", "р": "R", "с": "S", "т": "T", "у": "U", "ф": "F", "х": "H", "ц": "TS",
"ч": "CH", "ш": "SH", "щ": "SCH", "ъ": "", "ы": "Y", "ь": "", "э": "E", "ю": "YU",
"я": "YA",
}
out = []
for ch in text.lower():
if ch in table:
out.append(table[ch])
elif ch.isalnum() and ch.isascii():
out.append(ch.upper())
else:
out.append("_")
key = "_".join(filter(None, "".join(out).split("_")))
return ("CLICK_" + key)[:40] if key else "CLICK_UNKNOWN"
def unique_key(key, used):
"""
Keep generated keys distinct within one screen. Two "Все" links or two cards
whose captions collapse to the same slug would otherwise share a key, and a
shared key silently attributes one control's metrics to another.
"""
if key not in used:
used.add(key)
return key
for suffix in range(2, 100):
tail = "_" + str(suffix)
# Trim the stem BEFORE appending. Truncating afterwards chops the suffix off
# again, so two long captions keep colliding on the same 40-char key.
candidate = key[: 40 - len(tail)] + tail
if candidate not in used:
used.add(candidate)
return candidate
return key
def classify(text):
"""
Returns (eventKey, eventNameRu, targetScreenId, confidence).
Confidence "rule" means the key came from the table above - a naming convention
this project chose, NOT a key observed coming out of the app. Verified on the
live app, the real keys look like HOMEPAGEPAYMENTS and OPENWINDOWPAYMENT, so a
rule-derived key will not join against ClickHouse until it has been confirmed
with tools/monitor_live_events.py. Only "observed" means the key is real.
"""
low = text.lower()
for needles, key, name_ru, target in EVENT_RULES:
if any(n in low for n in needles):
return key, name_ru, target, "rule"
return slugify_event_key(text), "Нажатие «" + text[:40] + "»", None, "guessed"
def classify_native(element):
"""Native chrome hotspots, keyed off resource-id which is stable across releases."""
rid = element.get("resourceId", "")
label = element.get("contentDesc") or element.get("text") or ""
table = {
"toolbarAvatarImageView": ("PROFILE_ICON_CLICK", "Переход в Профиль", "profile_screen", "Аватар профиля"),
"action_telecomkz_account": ("TAB_CABINET_CLICK", "Вкладка «Кабинет»", "main_dashboard", "Вкладка «Кабинет»"),
"action_tv_plus": ("TAB_TV_PLUS_CLICK", "Вкладка «TV+»", None, "Вкладка «TV+»"),
"action_music": ("TAB_MUSIC_CLICK", "Вкладка «Музыка»", "music_screen", "Вкладка «Музыка»"),
"action_chats": ("TAB_CHATS_CLICK", "Вкладка «Чаты»", None, "Вкладка «Чаты»"),
"action_b2b": ("TAB_BUSINESS_CLICK", "Вкладка «Бизнес»", None, "Вкладка «Бизнес»"),
}
short = rid.split("/")[-1] if rid else ""
if short in table:
key, name_ru, target, label_ru = table[short]
return key, name_ru, target, label_ru, "rule"
desc = (element.get("contentDesc") or "").lower()
if "notification" in desc or "уведомл" in desc:
return "NOTIFICATIONS_OPEN", "Открытие уведомлений", None, "Уведомления", "rule"
if "navbar" in desc or "меню" in desc or "menu" in desc:
return "MENUCLICKED", "Нажатие на Меню", "side_menu", "Меню (правый навбар)", "rule"
key, name_ru, target, conf = classify(label or short or "native")
return key, name_ru, target, (label or short or "Нативный элемент"), conf
def wait_until_stable(session, timeout=20.0, quiet_polls=3, interval=0.6):
"""
Block until the page stops changing, or `timeout` elapses.
Returns (stable, seconds_waited). "Stable" means the fingerprint repeated
`quiet_polls` times in a row with no skeleton/loader elements left. Callers get
the flag so a capture taken from a still-moving page can be reported as such
rather than quietly saved as if it were finished.
"""
last = None
repeats = 0
started = time.time()
while time.time() - started < timeout:
try:
current = json.loads(session.evaluate(PAGE_FINGERPRINT_JS))
except T.DeviceError:
return False, time.time() - started
fingerprint = (current["n"], current["h"], current["t"])
if fingerprint == last and current["skeleton"] == 0:
repeats += 1
if repeats >= quiet_polls:
return True, time.time() - started
else:
repeats = 0
last = fingerprint
time.sleep(interval)
return False, time.time() - started
def capture_stitched(session, device, layout, settle=0.45, no_scroll=False):
"""
Scroll the page from top to bottom, capturing a native frame at each stop, and
compose them onto one canvas positioned by the measured scroll offset.
Returns (PIL image, geometry dict).
"""
screen_w = layout["screenWidth"]
screen_h = layout["screenHeight"]
wv_top = layout["webviewTop"]
wv_bottom = layout["webviewBottom"]
nav_top = layout["bottomNavTop"]
band_h = wv_bottom - wv_top
if not no_scroll:
session.evaluate("window.scrollTo(0, 0)")
time.sleep(settle)
geo = json.loads(session.evaluate("JSON.stringify(" + PAGE_GEOMETRY_JS + ")"))
scale = screen_w / max(geo["innerWidth"], 1)
# An overlay (side drawer, modal) covers a document that still reports a
# scrollable height, and scrolling it dismisses the overlay - so for those the
# capture is a single frame and the content band is exactly one viewport.
max_scroll_css = 0 if no_scroll else max(geo["maxScrollY"], 0)
content_h = round(max_scroll_css * scale) + band_h
first_png = T.screencap(device)
first = Image.open(io.BytesIO(first_png)).convert("RGB")
top_bar = first.crop((0, 0, screen_w, wv_top))
bottom_nav = first.crop((0, nav_top, screen_w, screen_h))
nav_h = bottom_nav.height
canvas = Image.new("RGB", (screen_w, content_h), (10, 15, 29))
canvas.paste(first.crop((0, wv_top, screen_w, wv_bottom)), (0, 0))
# Overlap each step slightly so a rounding error can never open a seam.
step_css = max(geo["innerHeight"] - 24, 40)
scroll_stops = []
pos = step_css
while pos < max_scroll_css - 1:
scroll_stops.append(pos)
pos += step_css
if max_scroll_css > 1:
scroll_stops.append(max_scroll_css)
# A fixed overlay (side menu, modal) sits on top of a document that still reports
# a scrollable height. window.scrollY then changes while the visible pixels do
# not, and pasting the next frame lower down duplicates the whole UI. So compare
# each frame against the last and stop as soon as the screen stops moving.
import numpy as np
prev_band = np.asarray(canvas.crop((0, 0, screen_w, band_h)), dtype=np.int16)
effective_content_h = content_h
for target_scroll in scroll_stops:
session.evaluate("window.scrollTo(0, " + str(target_scroll) + ")")
time.sleep(settle)
actual = session.evaluate("window.scrollY") or 0
frame = Image.open(io.BytesIO(T.screencap(device))).convert("RGB")
band = frame.crop((0, wv_top, screen_w, wv_bottom))
band_arr = np.asarray(band, dtype=np.int16)
if np.abs(band_arr - prev_band).mean() < 1.0:
effective_content_h = band_h
break
canvas.paste(band, (0, round(actual * scale)))
prev_band = band_arr
if effective_content_h != content_h:
canvas = canvas.crop((0, 0, screen_w, effective_content_h))
content_h = effective_content_h
if not no_scroll:
session.evaluate("window.scrollTo(0, 0)")
time.sleep(settle)
master = Image.new("RGB", (screen_w, wv_top + content_h + nav_h), (10, 15, 29))
master.paste(top_bar, (0, 0))
master.paste(canvas, (0, wv_top))
master.paste(bottom_nav, (0, wv_top + content_h))
geo.update(
{
"scale": scale,
"webviewTop": wv_top,
"contentHeight": content_h,
"navHeight": nav_h,
"totalHeight": master.height,
"scrollStops": len(scroll_stops) + 1,
}
)
return master, geo
def build_hotspots(dom_elements, layout, geo, catalog, source, viewport_only=False):
"""DOM elements + native chrome -> hotspots in stitched-screen space."""
scale = geo["scale"]
wv_top = geo["webviewTop"]
screen_w = layout["screenWidth"]
total_h = geo["totalHeight"]
nav_top_in_image = wv_top + geo["contentHeight"]
hotspots = []
seen_boxes = set()
used_keys = set()
# In no-scroll mode the image is exactly one viewport, so anything laid out below
# it is not in the picture and must not become a hotspot.
viewport_limit = wv_top + (layout["webviewBottom"] - wv_top) if viewport_only else None
for el in dom_elements:
x = round(el["cssLeft"] * scale)
y = wv_top + round(el["cssTop"] * scale)
w = round(el["cssWidth"] * scale)
h = round(el["cssHeight"] * scale)
# Horizontal carousels report boxes that run past the right edge of the
# screen. Clip to the visible canvas and drop anything entirely outside it.
x0, y0 = max(x, 0), max(y, 0)
x1, y1 = min(x + w, screen_w), min(y + h, total_h)
if viewport_limit is not None and y0 >= viewport_limit:
continue
if x1 - x0 < 20 or y1 - y0 < 15:
continue
box = (x0 // 8, y0 // 8, (x1 - x0) // 8, (y1 - y0) // 8)
if box in seen_boxes:
continue
seen_boxes.add(box)
key, name_ru, target, confidence = classify(el["text"])
if el.get("dataEvent"):
key, confidence = el["dataEvent"], "dom-attribute"
elif key in used_keys:
# Two different controls matched the same rule ("Мой автоплатеж" and
# "История платежей" both contain "платеж"). One key cannot describe both,
# and a duplicate silently joins the wrong ClickHouse rows to a button, so
# the later one falls back to a caption slug flagged as a guess.
key, name_ru, target, confidence = (
slugify_event_key(el["text"]),
"Нажатие «" + el["text"][:40] + "»",
None,
"guessed",
)
key = unique_key(key, used_keys)
hotspots.append(
T.attach_metrics(
{
"id": "hs_dom_" + str(len(hotspots) + 1),
"label": el["text"][:60],
"rect": {"x": x0, "y": y0, "width": x1 - x0, "height": y1 - y0},
"eventKey": key,
"eventNameRu": name_ru,
"category": "action",
"targetScreenId": target,
"source": "webview-dom",
"keyConfidence": confidence,
},
catalog,
source,
)
)
for i, el in enumerate(layout.get("nativeElements", [])):
rect = dict(el["rect"])
# Native chrome is fixed on screen: the top bar keeps its own coordinates and
# the bottom nav moves to wherever the nav band ended up in the tall image.
if rect["y"] >= layout["bottomNavTop"] - 4:
rect["y"] = nav_top_in_image + (rect["y"] - layout["bottomNavTop"])
key, name_ru, target, label_ru, confidence = classify_native(el)
hotspots.append(
T.attach_metrics(
{
"id": "hs_native_" + (el.get("resourceId", "").split("/")[-1] or str(i)),
"label": label_ru,
"rect": rect,
"eventKey": key,
"eventNameRu": name_ru,
"category": "navigation",
"targetScreenId": target,
"source": "native-uiautomator",
"keyConfidence": confidence,
},
catalog,
source,
)
)
return hotspots
def capture_native(screen_id, screen_name=None, category="Основное"):
"""
Capture a screen that has no WebView at all (Музыка, Чаты and the other native
tabs). Everything comes from uiautomator: one screencap plus every clickable
control on it. No scrolling - the accessibility tree only describes what is
currently rendered.
"""
T.validate_screen_id(screen_id)
device = T.get_device()
T.require_awake(device)
T.bring_app_to_front(device)
catalog = T.load_metrics_catalog()
source = T.metrics_source()
image = Image.open(io.BytesIO(T.screencap(device))).convert("RGB")
elements = T.collect_native_elements(device)
hotspots = []
used_keys = set()
for i, el in enumerate(elements):
label = el.get("label") or el.get("contentDesc") or el.get("text") or ""
short_rid = el.get("resourceId", "").split("/")[-1]
if not label and not short_rid:
continue
key, name_ru, target, label_ru, confidence = classify_native(el)
if confidence != "rule" or key in used_keys:
key = slugify_event_key(label or short_rid)
name_ru = "Нажатие «" + (label or short_rid)[:40] + "»"
target = None
confidence = "guessed"
key = unique_key(key, used_keys)
hotspots.append(
T.attach_metrics(
{
"id": "hs_native_" + (short_rid or str(i)),
"label": (label or label_ru or short_rid)[:60],
"rect": el["rect"],
"eventKey": key,
"eventNameRu": name_ru,
"category": "navigation",
"targetScreenId": target,
"source": "native-uiautomator",
"keyConfidence": confidence,
},
catalog,
source,
)
)
T.SCREENS_DIR.mkdir(parents=True, exist_ok=True)
image_path = T.SCREENS_DIR / (screen_id + "_long.png")
image.save(image_path, "PNG")
screen = {
"id": screen_id,
"name": screen_name or ("Экран " + screen_id),
"category": category,
"image": "/assets/screens/" + screen_id + "_long.png",
"isScrollable": False,
"viewportWidth": image.width,
"viewportHeight": image.height,
"totalHeight": image.height,
"surface": "native",
"capturedAt": int(time.time() * 1000),
# uiautomator only ever describes what is on screen now, so there is no
# renderer to wait for here.
"renderSettled": True,
"hotspots": hotspots,
}
app_map = T.load_app_map()
T.upsert_screen(app_map, screen)
app_map["metricsSource"] = source
T.save_app_map(app_map)
result = {
"success": True,
"screenId": screen_id,
"imageUrl": screen["image"],
"dimensions": {"width": image.width, "height": image.height},
"isLong": False,
"surface": "native",
"metricsSource": source,
"hotspots": hotspots,
"screen": screen,
"timestamp": int(time.time() * 1000),
}
with open(T.SCREENS_DIR / (screen_id + "_result.json"), "w", encoding="utf-8") as f:
json.dump(result, f, ensure_ascii=False, indent=2)
return result
def capture(
screen_id,
screen_name=None,
category="Основное",
settle=0.45,
no_scroll=False,
wait=20.0,
):
T.validate_screen_id(screen_id)
device, target = T.connect()
catalog = T.load_metrics_catalog()
source = T.metrics_source()
layout = T.probe_native_layout(device)
with T.CdpSession(target["webSocketDebuggerUrl"]) as session:
stable, waited = (True, 0.0)
if wait > 0:
stable, waited = wait_until_stable(session, timeout=wait)
if not stable:
print(
"warning: "
+ screen_id
+ " was still rendering after "
+ str(round(waited, 1))
+ "s; capturing anyway",
file=sys.stderr,
)
image, geo = capture_stitched(session, device, layout, settle=settle, no_scroll=no_scroll)
dom = json.loads(
session.evaluate("JSON.stringify(" + dom_elements_js(no_scroll) + ")")
)
hotspots = build_hotspots(
dom.get("elements", []), layout, geo, catalog, source, viewport_only=no_scroll
)
T.SCREENS_DIR.mkdir(parents=True, exist_ok=True)
image_path = T.SCREENS_DIR / (screen_id + "_long.png")
image.save(image_path, "PNG")
screen = {
"id": screen_id,
"name": screen_name or ("Экран " + screen_id),
"category": category,
"image": "/assets/screens/" + screen_id + "_long.png",
"isScrollable": image.height > layout["screenHeight"],
"viewportWidth": layout["screenWidth"],
"viewportHeight": layout["screenHeight"],
"totalHeight": image.height,
"route": geo.get("route"),
"capturedAt": int(time.time() * 1000),
# False means the page was still painting when this was taken - the capture
# is kept, but flagged so it is obvious the screen may be incomplete.
"renderSettled": stable,
"hotspots": hotspots,
}
app_map = T.load_app_map()
T.upsert_screen(app_map, screen)
app_map["metricsSource"] = source
T.save_app_map(app_map)
result = {
"success": True,
"screenId": screen_id,
"imageUrl": screen["image"],
"dimensions": {"width": image.width, "height": image.height},
"isLong": screen["isScrollable"],
"route": geo.get("route"),
"scrollStops": geo.get("scrollStops"),
"metricsSource": source,
"hotspots": hotspots,
"screen": screen,
"timestamp": int(time.time() * 1000),
}
with open(T.SCREENS_DIR / (screen_id + "_result.json"), "w", encoding="utf-8") as f:
json.dump(result, f, ensure_ascii=False, indent=2)
return result
def main():
parser = argparse.ArgumentParser(description="Capture a TelecomKz screen and map its hotspots.")
parser.add_argument("screen_id", nargs="?", default="main_dashboard")
parser.add_argument("--name", default=None, help="Human readable screen name")
parser.add_argument("--category", default="Основное")
parser.add_argument("--settle", type=float, default=0.45, help="Seconds to wait after each scroll")
parser.add_argument(
"--wait",
type=float,
default=20.0,
help="Seconds to wait for the page to stop rendering before capturing",
)
parser.add_argument(
"--native",
action="store_true",
help="Screen has no WebView: map it from uiautomator instead of the DOM",
)
parser.add_argument(
"--no-scroll",
action="store_true",
help="Capture a single viewport without scrolling (side drawer, modals)",
)
args = parser.parse_args()
try:
if args.native:
result = capture_native(args.screen_id, args.name, args.category)
else:
result = capture(
args.screen_id,
args.name,
args.category,
args.settle,
args.no_scroll,
args.wait,
)
except (T.DeviceError, ValueError) as exc:
print(json.dumps({"success": False, "error": str(exc)}, ensure_ascii=False))
return 1
# stdout is the API contract with the Vite middleware: one JSON object.
print(
json.dumps(
{k: v for k, v in result.items() if k not in ("hotspots", "screen")},
ensure_ascii=False,
)
)
print(
"Captured "
+ args.screen_id
+ ": "
+ str(result["dimensions"]["width"])
+ "x"
+ str(result["dimensions"]["height"])
+ ", "
+ str(len(result["hotspots"]))
+ " hotspots",
file=sys.stderr,
)
return 0
if __name__ == "__main__":
sys.exit(main())