Subversion Repositories SmartDukaan

Rev

Show changed files | Details | Compare with Previous | Blame | RSS feed

Filtering Options

Rev Age Author Path Log message Diff
37565 27 d 7 h amit /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/ IMEI activation: 20s tick lanes for oppo/realme/vivo, idle once the day clears

Replaces the two once-a-day passes with chunk-per-tick lanes. Each tick takes one
chunk and returns; when a brand's pool comes back empty its turn is skipped, and
when everything is clear the ticks do nothing and stay silent until midnight.

- browser lane (StandAlone): one chunk of 25 for one brand every 20s, oppo and
realme by turns. Still exactly one ChromeDriver alive at a time.
- vivo lane: its own tick, chunk of 250. It must not share the browser lane --
0.24s an imei against oppo's 10.2s means it would need ~70 hours behind them
for work it does alone in 39 minutes.

The snapshot is gone; the pool query is the cursor. That is what makes a restart
cost one chunk instead of the day: the 12:03 restart on 09-Sep forfeited ~4,000
lookups and the whole afternoon, and last_finish had read -1 for three days.

The snapshot existed to stop the re-ask loop (realme, 29-Aug: 4,524 requests
against 1,004 distinct imeis). That is now closed at the source instead -- oppo,
realme and motorola stamp every imei they asked about, not just the ones that
produced a map entry, so a failed lookup rests until tomorrow rather than coming
back on the next tick. Motorola is fixed pre-emptively; nothing schedules it yet.

Also:
- secondary and tertiary are merged by turns rather than concatenated. Safe while
a pass walked to the end; without that guarantee oppo's 163 tertiary serials sat
behind 3,819 secondary ones and would only be reached on a day that cleared.
- vivo abandons a tick rather than the chunk when the captcha solver returns no
code -- one probe per 20s while it is down instead of 250, and no rows rested
over a transient outage.
- the funnel gauges move from a pass to a day. due is measured on the first tick
after midnight, the rest accumulate, and last_finish_epoch becomes a real
completion clock. Truncation is detected at the midnight rollover, which is the
case that never reaches an end-of-run at all.

This does not create capacity. At the 21s/imei measured on 09-Sep the pool still
needs ~35 hours and will not clear; it now rolls over visibly instead of silently.
The lever for that is DAYS=1.
 
37472 36 d 18 h amit /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/ IMEI activation: one snapshotted daily pass per brand, on one thread, with per-brand metrics

Four @Scheduled jobs every 5 minutes become two daily passes. Oppo and realme
share one thread and alternate in 25-imei chunks, so exactly one ChromeDriver is
alive at a time instead of four; vivo keeps its own thread since it is direct
HTTP and does not contend for a browser.

The pass snapshots its pool before any browser starts and walks that list to the
end. It never re-queries, and that is the actual fix. A failed lookup never
reaches dateMap.put, so no row is written, so createTimestamp is not bumped, so
the imei was eligible again on the next tick five minutes later. Measured 29-Aug:
realme issued 4,524 requests against 1,004 distinct imeis -- 4.5 asks each, 78%
of the day's budget spent re-asking -- while oppo, which rarely fails, sat at
1.03. More requests hardened the block, which caused more failures. A pass bounds
that: a failure costs one retry tomorrow, never one in five minutes.

This supersedes the r37447/r37448/r37449 argument about driver count, which was
about the wrong variable. That argument blamed realme's collapse on CPU
contention pushing the captcha render past the element waits. The logs do not
support it: on 29-Aug oppo took ZERO canvas timeouts across all 24 hours on the
same box, same six cores, same driver count, same captcha vendor, load average
0.9 -- including the 15:00-23:00 window in which realme solved nothing at all.
Realme's own canvas wait is 15s against oppo's 8s, so the longer wait is the one
expiring. What realme's timeout rate tracks is its own daily request volume, and
it resets at midnight: 920/day -> 0.3%, 3,467/day -> 28%, 4,524/day -> 75%. That
is realme.com declining to serve the widget.

DAYS=0 is deliberate and is not an off-by-one: the pool filter is
createTimestamp < now().atStartOfDay().minusDays(DAYS), so DAYS=1 measures
against yesterday midnight and silently yields a two-day cadence, which is what
oppo and realme were running.

Sizing measured on prod for a midnight start: oppo 4,133 and realme 2,118 imeis,
11.7h + 8.4h = 20.1 hours of a single thread. It fits with no slack; if the
'pass finished' counts come in short of 'pass starting', the lever is DAYS=1
rather than a second thread.

Observability: ImeiActivationGauges publishes the funnel per brand on
/actuator/prometheus, which alloy already scrapes on this host -- due, churned,
captcha_shown, captcha_solved, answered, dates_found, errors, run_seconds and
last_finish_epoch. Each stage fails differently and says what broke. Rates are
left to PromQL. The stage that matters for health is answered: churned>0 with
answered==0 is precisely the shape of both silent outages this year (oppo wrote
nothing for a week; the vivo captcha solver was dead for 46 days). dates_found is
deliberately NOT a health signal -- when the multi-year backlog drained at the
end of August, yield fell from ~100% to 2-3% on the same day across all three
brands with nothing broken.

Nagios cleanup: the Nagios server and every NRPE daemon are gone, so
WriteToPropertiesFile and the commented-out blocks that fed
nagios-cron.properties are deleted, and NagiosMonitorTasks is renamed
BalanceMonitorTasks for the transport it actually uses. Noted there that nothing
calls it -- there is no @Scheduled entry and no other caller -- which is why both
balance gauges have always read -1.