(root)/ – Rev 37678
Rev 37677 |
Rev 37679 |
Go to most recent revision |
Last modification |
Compare with Previous |
View Log
| RSS feed
Last modification
- Rev 37678 2026-09-17 15:55:49
- Author: amit
- Log message:
- Replace the knowlarity insights chrome scrape with SR's own JSON API
Drops the last unattended headless chrome out of the fofo tomcat. The scheduled
insights job ran 8 times a day and each run started an ~850MB chrome tree inside
the tomcat JVM's host -- a box holding -Xmx8g tomcat plus a -Xmx2g cron jar on
16GB that has been kernel-OOM-killed twice with tomcat the victim.
It was also losing data the whole time. Verified on prod across 7 consecutive
days: every row of cs.agent_daily_insight has logged_in_seconds, break_seconds,
available_seconds, talk_seconds, calls_answered, missed_calls and total_calls
set to 0. Two separate causes, both fixed here:
- parseTimeToSeconds split on ':' expecting 'HH:mm:ss', but the table renders
'2h 50m 39s'. A one-element split fell through to return 0, and because it
never reached the NumberFormatException branch it did not even warn. It now
parses the h/m/s shape and still accepts HH:mm:ss and HH:mm.
- the scrape never captured the call counts at all. The API carries them.
The auth chain is four calls and is not guessable, so it is documented in
KnowlarityApiClient: POST /vr/sr_login/ establishes the session and returns an
HS256 token that the API REJECTS; GET /newsr/user_details yields new_sr_ui_url
carrying a one-shot SSO blob (in a browser this hop is javascript, so it is
invisible to anything that merely follows redirects); POST /vr/sso_login/
exchanges that blob for the RS256 token the API accepts, whose claims embed the
srsessionid and so bind it to the session; GET /newsr/agents_insights/ with
header jwtAuthorization. Wrong token and wrong header name both answer
'Invalid token', so the error never tells you which mistake you made. The window
parameters are start_time/end_time -- start_date/end_date authenticates fine and
returns 'Error in API'.
Redirects are followed BY HAND. setInstanceFollowRedirects(true) exposes only
the final response's headers, so the Set-Cookie issued on the intermediate hops
is lost, the session never forms and user_details answers with an HTML error
page. This cost a debugging cycle; the reason is commented at the call site.
Verified against the live account before committing: 13 agents returned,
calls_offered == calls_answered + missed_calls holds for all 13, durations match
the previous scrape to within the elapsed window (~55s), and formatSeconds /
parseTimeToSeconds round-trip cleanly over the real values.
DTO fields stay the same display strings ('2h 50m 39s'), so every existing
reader is unaffected; only the previously-zero numeric columns change.
KnowlarityScraperService still uses ChromeDriver for operator-triggered break-log
scrapes, so the selenium dependency stays for now. Its @Scheduled annotations are
already commented out, so nothing unattended starts a browser any more.