Subversion Repositories SmartDukaan

Rev

Go to most recent revision | Show changed files | Details | Compare with Previous | Blame | RSS feed

Filtering Options

Rev Age Author Path Log message Diff
37472 37 d 7 h amit /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/ IMEI activation: one snapshotted daily pass per brand, on one thread, with per-brand metrics

Four @Scheduled jobs every 5 minutes become two daily passes. Oppo and realme
share one thread and alternate in 25-imei chunks, so exactly one ChromeDriver is
alive at a time instead of four; vivo keeps its own thread since it is direct
HTTP and does not contend for a browser.

The pass snapshots its pool before any browser starts and walks that list to the
end. It never re-queries, and that is the actual fix. A failed lookup never
reaches dateMap.put, so no row is written, so createTimestamp is not bumped, so
the imei was eligible again on the next tick five minutes later. Measured 29-Aug:
realme issued 4,524 requests against 1,004 distinct imeis -- 4.5 asks each, 78%
of the day's budget spent re-asking -- while oppo, which rarely fails, sat at
1.03. More requests hardened the block, which caused more failures. A pass bounds
that: a failure costs one retry tomorrow, never one in five minutes.

This supersedes the r37447/r37448/r37449 argument about driver count, which was
about the wrong variable. That argument blamed realme's collapse on CPU
contention pushing the captcha render past the element waits. The logs do not
support it: on 29-Aug oppo took ZERO canvas timeouts across all 24 hours on the
same box, same six cores, same driver count, same captcha vendor, load average
0.9 -- including the 15:00-23:00 window in which realme solved nothing at all.
Realme's own canvas wait is 15s against oppo's 8s, so the longer wait is the one
expiring. What realme's timeout rate tracks is its own daily request volume, and
it resets at midnight: 920/day -> 0.3%, 3,467/day -> 28%, 4,524/day -> 75%. That
is realme.com declining to serve the widget.

DAYS=0 is deliberate and is not an off-by-one: the pool filter is
createTimestamp < now().atStartOfDay().minusDays(DAYS), so DAYS=1 measures
against yesterday midnight and silently yields a two-day cadence, which is what
oppo and realme were running.

Sizing measured on prod for a midnight start: oppo 4,133 and realme 2,118 imeis,
11.7h + 8.4h = 20.1 hours of a single thread. It fits with no slack; if the
'pass finished' counts come in short of 'pass starting', the lever is DAYS=1
rather than a second thread.

Observability: ImeiActivationGauges publishes the funnel per brand on
/actuator/prometheus, which alloy already scrapes on this host -- due, churned,
captcha_shown, captcha_solved, answered, dates_found, errors, run_seconds and
last_finish_epoch. Each stage fails differently and says what broke. Rates are
left to PromQL. The stage that matters for health is answered: churned>0 with
answered==0 is precisely the shape of both silent outages this year (oppo wrote
nothing for a week; the vivo captcha solver was dead for 46 days). dates_found is
deliberately NOT a health signal -- when the multi-year backlog drained at the
end of August, yield fell from ~100% to 2-3% on the same day across all three
brands with nothing broken.

Nagios cleanup: the Nagios server and every NRPE daemon are gone, so
WriteToPropertiesFile and the commented-out blocks that fed
nagios-cron.properties are deleted, and NagiosMonitorTasks is renamed
BalanceMonitorTasks for the transport it actually uses. Noted there that nothing
calls it -- there is no @Scheduled entry and no other caller -- which is why both
balance gauges have always read -1.
 
37448 40 d 11 h amit /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/scheduled/ Daily re-check for all brands; Oppo back to parallel pools

Two changes.

1. Re-check window 4 days (secondary) and 2 days (tertiary) -> 1 day everywhere.
Daily demand becomes the full universe rather than a fraction of it:

Oppo 4,798 + 4,040 = 8,838/day
Vivo 9,407 + 675 = 10,082/day (doing 15,345 -- fine)
Realme 1,973 + 1,081 = 3,054/day

2. Oppo's two pools run in PARALLEL again, reverting the merge in r37447 for that
brand only. Realme stays merged.

The merge was a straight trade of throughput for memory and oppo could not
afford it. Measured over 32 minutes and again over an hour the next morning:
3,555 then 3,456/day against 5,280 before merging. Batch cadence settled at a
very regular ~15.5 min per cycle, so a 30-imei merged batch takes ~10.5 min =
~21s/imei, against the 10.2s it managed unmerged. At 21s the ceiling is
86400/21 = 4,114/day even with zero idle, so no batch size and no shorter
fixedDelay could have reached 8,838. Serialising simply costs more per imei
here than running two browsers does.

Realme keeps the merge: it needs 3,054/day and delivers 2,952 merged, so a
small size bump covers it without a second browser.

Sizes: oppo 25 per pool (2 jobs in parallel), realme 12+12 merged, vivo 50+10
unchanged. Vivo has already cleared its entire secondary backlog -- the pool
reads 0 and both lists come back empty -- which is what the batch of 50 was for.

Cost: oppo goes back to two concurrent drivers, so the fleet is 3 rather than 2,
roughly +700MB. Acceptable against the ~2GB freed today by reaping orphaned
browsers and capping retries, but it is the reason realme was left merged.

Sizes are a starting point, not a final answer: oppo's per-imei time differs
markedly between merged and parallel modes, so re-measure before tuning further.
 
37446 41 d 3 h amit /trunk/ Per-brand batch sizes, and stop chrome forking a GPU process it cannot use

maxResults was hardcoded in the shared repository methods, so Oppo and Vivo were
forced to the same secondary batch (10) and all three to the same tertiary (10).
It is now a parameter, set per brand at the call site.

Sizing is arithmetic, from measured IN-BATCH per-imei time. Solving
M * 86400 / (300 + M*t) = needed/day:

brand needed/day t M required set to
Vivo 9,407 0.8s 36 50 clears, ~12,700/day
Realme 1,973 13.4s 10 10 was 5 = ~1,177/day, short
Oppo 4,798 14.6s 88 10 HELD, see below

Correcting an earlier measurement of mine: I reported Vivo at 13.6s per imei and
concluded its backlog could not be cleared. That averaged across the ~300s idle
gaps BETWEEN batches. In-batch it is 0.8s -- Vivo is 17x faster than I said, is
idle ~97% of the time, and 50 clears its pool comfortably. There is no wait in
the Vivo path; it is simply fast.

Oppo is deliberately NOT raised. At 14.6s it would need M=88, which means
20-minute batches and near-permanent chrome sessions. But that 14.6s predates the
retry cap (r37445), which cuts exhausted imeis from 20 attempts to 7 and should
drop it sharply. Re-measure before sizing Oppo, rather than guessing high on a
box with 3GB free.

Also: --disable-gpu, --disable-dev-shm-usage, --disable-software-rasterizer on
both selenium tasks. Headless needs no GPU yet chrome forks a gpu-process per
browser -- 6 were alive across the fleet, pure overhead. No behaviour change.

Batch size does not raise peak concurrency (fixedDelay means one batch per job at
a time, so never more than 4 drivers). It raises DUTY CYCLE, which converts
chrome's footprint from intermittent to sustained. That matters here: tomcat is
9.4GB, available is ~3GB, and the two OOM kills this month both took tomcat.

Cron-only deploy. The dao signature change has no callers outside cron.
 
37408 42 d 11 h amit /trunk/ Vivo IMEI activation: stop spending captchas on line items with no IMEI

76% of production captcha rejections were line items whose serial number is
null. Vivo answers those with {"msg":"参数为空"} -- "parameter is empty" --
and status 0, which this code recorded as a captcha failure. So a correctly
solved captcha looked wrong, no activated_imei row was written, the line item
stayed pending, and it came back every 5 minutes indefinitely. Those rows were
permanently consuming roughly 43% of the run quota, which is why every tick ran
full at 20/20 and the backlog never drained.

It also made the model look far worse than it is: measured accept rate 44%,
while the same solver scores 85-93% when the IMEI is present. True captcha
accuracy is around 77%.

Fixed at source: both named queries now exclude null and blank serial numbers,
so such line items never enter the pool (this also covers the Realme caller).
The loop additionally skips them before fetching a captcha, so no captcha,
solver call or Vivo request is spent discovering it.

Found via the diagnostics added in r37407 -- the previous log line recorded
five words and discarded the response that named the cause.
 
37407 42 d 11 h amit /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/scheduled/ Vivo IMEI activation: log why a captcha was rejected

Production refuses ~56% of submissions, but reading the stored images by hand
showed 9 of 10 carried the CORRECT code. Replicating the request from the same
box gives 85-93%, and every external difference was measured and ruled out:
real vs dummy imei (7/8 both), user-agent (83% both), session reuse (74 vs 75%),
two concurrent jobs (80% both), source IP, and fetch-to-submit delay from 0 to
10s (90-91% throughout). Vivo's reply is byte-identical in every failure, so the
payload carries no discriminator.

That leaves something inside this process, and the old log line recorded five
words and discarded the evidence. It now captures the code submitted, the imei,
the real fetch-to-submit latency, the session cookies actually held at submit
time, and Vivo's full response.

The cookies matter most: if the SESSION cookie is being dropped or rotated
between fetching the captcha and submitting it, the captcha would be validated
against the wrong session and refused despite a correct code -- which fits every
measurement above and is invisible from outside the JVM.
 
37402 43 d 2 h amit /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/scheduled/ Vivo IMEI activation: report Vivo's captcha verdict back to the solver

Only this cron ever learns whether Vivo accepted a captcha (status 0 means it
was wrong), so it is the only place that can label the solver's training data.
After each checkCode response it now posts the image, the code and the outcome
to the solver's /verdict endpoint, which files the sample as accepted (the
prediction was right - a free label) or rejected (needs a human to label it).

The solver has been running at a measured 22.3% accept rate and no training
data was ever collected, so the model could not be improved at all.

Strictly best-effort: 2s connect / 3s read timeouts and every exception
swallowed. Data collection must never slow down or break IMEI activation.

The shared secret is read from CAPTCHA_VERDICT_TOKEN in the environment,
falling back to a captcha.verdict.token property, so it need not be committed.
Without a token the call is skipped entirely.

Also drops /tmp/captcha.jpg: the captcha bytes are needed in memory for the
verdict report anyway, so the two concurrently-scheduled Vivo jobs no longer
share one file and overwrite each other's image.
 
37393 43 d 6 h amit /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/scheduled/ Vivo IMEI activation: validate captcha solver response before submitting to Vivo

The captcha solver at 45.79.121.178 was down from 09-Jul-2026 to 25-Aug-2026
(uwsgi never restarted after a host reboot). nginx returned a 502 page for
every request. RestClient.executeJson returns the response body regardless of
HTTP status, so CaptchaService handed that HTML page to Vivo as the captcha
code and Vivo logged it as 'Found invalid captcha' - indistinguishable from an
ordinary wrong guess. The outage went unnoticed for 46 days.

CaptchaService now checks the response against the solver model's own 31-class
alphabet (^[1-9A-HK-NP-Z]{4}$) and returns null for anything else, logging the
offending payload. VivoImeiActivationService skips the IMEI when the code is
null instead of posting the garbage; it is retried on the next run.

Also drops a System.out.println of the full base64 captcha image.
 
36580 143 d 8 h amit /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/scheduled/ Adjust IMEI activation deferral: secondary 4 days, tertiary 2 days for Vivo/Oppo/Realme  
36420 162 d 5 h amit /trunk/profitmandi-cron/ OkHttp→RestClient migration for IMEI activation services (Itel, Tecno, Vivo). Added test deps. Updated RunOnceTasks, ScheduledTasks, OrderTrackingService.  
36260 177 d 7 h amit /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/scheduled/ Fix: add @Transactional(NOT_SUPPORTED) on Vivo methods to suspend ScheduledTasks outer transaction  
36253 178 d 13 h amit /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/scheduled/ Separate secondary/tertiary IMEI activation crons for Vivo/Oppo/Realme, perf fixes: shared saveActivation, Response leak fixes, /tmp cleanup, OpenCV static init, early break, remove class-level @Transactional from StandAlone  
33453 844 d 8 h amit.gupta /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/scheduled/ Do not run Vivo Activations on 1st of every month  
31128 1433 d 7 h amit.gupta /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/scheduled/ Fixed Scheduling issues  
30937 1489 d 7 h amit.gupta /trunk/ Fixed activation logic  
30430 1603 d 12 h tejbeer /trunk/ change  
30376 1619 d 2 h amit.gupta /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/scheduled/ Added Einvoice Files  
30371 1619 d 5 h amit.gupta /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/scheduled/ Added Einvoice Files  
30369 1619 d 6 h amit.gupta /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/scheduled/ Added Einvoice Files  
30343 1625 d 5 h amit.gupta /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/scheduled/ Added Einvoice Files  
30337 1626 d 5 h amit.gupta /trunk/profitmandi-cron/src/main/java/com/smartdukaan/cron/scheduled/ Fixed ahead issue  

Show All