Subversion Repositories SmartDukaan

Rev

Show changed files | Details | Compare with Previous | Blame | RSS feed

Filtering Options

Rev Age Author Path Log message Diff
37788 10 d 3 h amit /trunk/profitmandi-cron/ activation: stop a broken JVM from spending the day's imei pool, and take the
selenium atom read out of the nested jar

Oppo/Realme activation collapsed on 21-22 Sep. Every WebDriver command failed with
java.util.zip.ZipException reading a Selenium JS atom:

W3CHttpCommandCodec.amendParameters:227 -> executeAtom:397
-> com.google.common.io.Resources.toString
-> org.springframework.boot.loader.jar.ZipInflaterInputStream.read -> ZipException

Scale, from fofo.activated_imei: a ~30-50 errors/day baseline became 17,508 on 21-Sep
and 27,768 on 22-Sep. On 22-Sep it touched 3,920 realme imeis for 0 dates and 4,824
oppo for 96, against a normal 76-100% hit rate. Realme burned its entire day pool by
11:07 and then correctly went quiet, having answered nothing.

TWO INDEPENDENT FAULTS, one fixed each way.

1. The pool was spent on an outage. restUnanswered stamps every unanswered imei so the
20-second tick advances instead of re-handing the same rows -- correct for a per-imei
failure, catastrophic for a systemic one, because a stamped row does not come back
until its rest expires. So a JVM that cannot read a jar quietly consumed a day of
payout data. Now: if NOTHING in the chunk was answered and the chunk had more than one
imei, that is infrastructure rather than a verdict, and nothing is stamped. Cost is a
re-ask of the same chunk next tick -- loud and self-limiting -- instead of the day.
Same class of bug as the carlcare transient refusal (r37537): far-end/our-end noise
must never be recorded as an answer.

2. The read itself. Selenium loads its atoms as classpath RESOURCES on essentially
every command, which inside a fat jar is a nested-jar read on the Spring Boot 2.0.2
(2018) loader. Three different inflater errors appeared on prod -- "invalid stored
block lengths", "invalid distance too far back", "invalid code lengths set" -- on a jar
whose outer AND extracted nested archives both pass `unzip -t`. Intact bytes with three
distinct inflater failures is a reader fault, not a file fault. bootJar now sets
requiresUnpack for selenium-remote-driver, so it is extracted to a real file at launch
and the atom read never touches the nested reader. Verified in the built jar: the entry
carries UNPACK:<sha1> and is STORED rather than DEFLATED.

TRAPS WORTH RECORDING.

- A restart is NOT a diagnosis here. It was restarted 13:16 on 22-Sep and still failed
for three more hours at ~1,260/hour, then a 16:36 restart came up clean -- same jar,
mtime unchanged. Anyone reading "restart fixed it" should distrust it.
- Concurrency alone does not explain it: up to 2 scheduler pools ran these tasks per
minute in the broken window AND in the healthy one.
- The GlitchTip board under-reported this badly (#1495 lastSeen 17-Sep while the log
held thousands on 21-22 Sep), so the board is not a reliable outage signal for cron.
Judge this lane on dates written in fofo.activated_imei, not on issue counts.
- Deploy the cron jar stop -> replace -> start. The jar on disk was overwritten in place
at 12:31 on 21-Sep while a JVM held it open, which is how this started.
 
37726 18 d 2 h amit /trunk/ sentry: stop developer laptops reporting to the live GlitchTip board

GlitchTip #586 was 57 events tagged environment=production whose stack read
/opt/homebrew/Cellar/tomcat@8/8.5.100/libexec/... with server_name set to a
developer's machine. Nothing was wrong on prod: a laptop was posting into the
production project and was indistinguishable from it.

Two things combined to allow that. The DSN lives in log4j2.xml, which ships inside
every build, so any machine running this code can report. And the Sentry SDK
defaults `environment` to "production" when it is not set -- which it never was --
so local runs arrived pre-labelled as prod.

Adds sentry.properties to each module, read off the classpath by the SDK itself
(io.sentry.config.PropertiesProviderFactory) and merged over the appender's config.
Both keys used here are honoured by io.sentry.ExternalOptions in 7.22.6 (verified
against the jar): `enabled` and `environment`.

The COMMITTED values are the safe ones -- enabled=false, environment=dev -- so a
plain local build is silent. build.gradle rewrites both from -Penv= alongside the
env.property it already writes, so only a deliberate -Penv=staging|prod build
reports, and it carries the right environment tag. tasks.build.doLast restores the
safe default afterwards, mirroring the existing handling of env.property.

Verified both directions: default build leaves enabled=false/environment=dev,
-Penv=prod yields enabled=true/environment=prod.

Note this makes the board trustworthy rather than merely quieter: events can now be
filtered on environment, and anything unlabelled is a build that predates this.
 
37491 36 d 6 h amit /trunk/ errors: ship ERROR events to GlitchTip via the log4j2 Sentry appender

io.sentry:sentry-log4j2:7.22.6 in all three deployables. Pinned to 7.x because
8.x drops Java 8; verified class major version 52.

minimumEventLevel=ERROR is load-bearing. It only became safe after r37490 moved
business validations to WARN -- 32% of the ERROR stream was HTTP 400s where the
user is simply told what to fix, and sending those would have made "insufficient
balance" the top issue and buried real bugs. Breadcrumbs come from INFO so an
issue arrives with the log lines that preceded it.

Attached to the application loggers only; framework noise is not our bug.

web's log4j2.xml is committed here. cron's and fofo's are held back: both carry
uncommitted local development paths (a user.home expansion, and an absolute
/Users path) that would break production logging and the Alloy log tailing.
Their Sentry blocks are staged locally and should land with whoever owns those
path edits.
 
37338 50 d 5 h amit /trunk/profitmandi-cron/ Remove offer-circular ingest from cron - it now runs in the portal

Counterpart to the fofo commit that moved the parse into profitmandi-fofo. Nothing is
lost: all 13 parser classes and the test moved verbatim, and CircularIngestScheduler
was replaced by an executor-driven runner in the portal.

- com.smartdukaan.cron.offercircular deleted, main and test
- tabula dependency removed; PDFBox no longer enters this artifact at all
- dumpCircularClasspath moved to fofo, where the parser and its classpath now live

This module no longer knows anything about offer circulars, so a cron rebuild is no
longer a prerequisite for the feature to work - which was the entire problem.
 
37334 50 d 6 h amit /trunk/profitmandi-cron/ Sweep stalled offer circulars before each ingest run, and drop stale scaffolding

- The scheduler now calls reclaimStalledProcessing before looking for new work, so
a circular stranded in PROCESSING by a dead or redeployed worker is failed with a
reason instead of staying invisible forever. Threshold is 30 minutes against a
parse that takes seconds, so it can only ever catch a genuinely dead worker. The
sweep is wrapped so it can never stop the actual ingest.
- dumpCircularClasspath was labelled a temporary helper for diffing against the
reference Python. That Python has been deleted, but the task is what lets an
ingest be reproduced and measured locally per scripts/offer_circular/README.md,
so the comment now says what it is for rather than telling the next reader to
delete it.
- CircularExtractor's javadoc claimed verification against the Python reference.
That proved equivalence, not correctness - both shared the missing-memory-unit
bug that bound offers to the wrong SKU. Reworded so it cannot be read as a
correctness guarantee, and points at the fixture and ProductNamesTest instead.

Verified: full ingest of the Aug'26 circular is byte-identical to the reference -
239 offers, 663 products, 279 AUTO_EXACT. ProductNamesTest 11/11.
 
37330 50 d 21 h amit /trunk/profitmandi-cron/ Add Pine Labs affordability circular parsing and ingest

Turns the monthly OEM "Mobile & Laptop Offers" PDF into the offers schema: tabula
extraction to nine verbatim columns per row, per-cell parsers for benefit, tenure,
bank and footnote text, catalog matching, and a scheduler that claims DRAFT
documents uploaded from the FOFO portal and emails the uploader the outcome.

Ships inert. offer.circular.ingest.enabled defaults to false, so the scheduler does
nothing until an environment opts in, and offer.circular.review.url defaults to
empty - neither key is required for the context to start.

- tabula added here and not in profitmandi-common so PDFBox never reaches the
web/fofo WARs; bouncycastle, slf4j-simple and jai-imageio excluded (version
clash, duplicate SLF4J binding, and unused image decoding respectively)
- offer_raw_row holds all nine columns verbatim, so re-parsing reads the table and
never the PDF again
- product_alias is consulted before matching, so a human confirmation recorded once
keeps applying every following month

ProductNamesTest locks in the product-name parsing, which decides which SKU an
offer's money lands on. Every case there is a real mis-parse, and they share one
root cause: characters or digits belonging to the model name being eaten as memory
or stripped as punctuation. Two worth naming:

- A memory spec with no GB/TB unit is still a memory spec. Motorola writes
"(8+256)" and vivo "(8+256G)"; unrecognised, the matcher believed no size was
given and bound the offer to an arbitrary sibling - a 1,000 Edge 60 Pro 8+256
offer and a 2,000 12+256 offer landed on the same SKU.
- Two variant groups written back-to-back are two products. "Edge 70 Pro
(8+256)(12+256)" stayed one product bound to a single SKU while the 12+256
variant silently got no offer at all.

Bundled accessories are deliberately NOT stripped back to the bare phone. Doing so
resolves 30 CATALOG_GAP rows and looks safe on Oppo, whose bundled and bare rows
carry identical values - but vivo caps X300 Pro(16+512G) at 10,000 on its own row
and 11,000 on the "+Extender" row, the difference being the Extender. Merging them
would let the bundle's cap be claimed on a phone sold without the accessory.
Whether a bundle offer transfers to the bare SKU is a commercial question the PDF
does not answer, so it is a manual coverage decision, not a parsing rule.

Verified against the Aug'26 circular on the local DB: 324 rows in the PDF, 239
loaded, producing 361 benefits, 511 tenures, 911 bank links and 663 products, of
which 74% resolve automatically. ProductNamesTest 11/11.
 
36420 162 d 0 h amit /trunk/profitmandi-cron/ OkHttp→RestClient migration for IMEI activation services (Itel, Tecno, Vivo). Added test deps. Updated RunOnceTasks, ScheduledTasks, OrderTrackingService.  
36020 203 d 0 h amit /trunk/profitmandi-cron/ Add Knowlarity call monitor cron scheduler - 10AM start, 10PM stop, 15min health check. Auto-login with username/password to fetch queue UUIDs.  
35851 224 d 0 h amit /trunk/profitmandi-cron/ Revert OpenCV from 4.9.0 to 3.4.2 - native libs not compatible with server environment  
35832 226 d 2 h amit /trunk/ Unify property loading: all modules use runtime profile with shared properties from dao, remove duplicated DB/Hibernate/HikariCP/integration keys from module files  
34860 426 d 23 h ranu /trunk/ razorpay x automate payment with rabbit mq  
34688 471 d 7 h amit.gupta /trunk/profitmandi-cron/ Fixed  
34554 511 d 3 h tejus.lohani /trunk/ cron job monitoring development related files  
34101 639 d 1 h amit.gupta /trunk/profitmandi-cron/ Added org.jetbrains  
34100 639 d 3 h amit.gupta /trunk/profitmandi-cron/ Added ok http  
34094 644 d 9 h amit.gupta /trunk/profitmandi-cron/ useless  
34039 667 d 1 h vikas.jangra /trunk/profitmandi-cron/ Resolved push notifications  
30362 1622 d 1 h amit.gupta /trunk/profitmandi-cron/ Added Einvoice Files  
30356 1623 d 0 h amit.gupta /trunk/profitmandi-cron/ Fixed ahead issue  
30212 1656 d 1 h amit.gupta /trunk/profitmandi-cron/ Fixed ahead issue  

Show All