← Ventures — GCStore Constellation
GCruiter — Level-2 Deep Dive
GCruiter is a candidate-facing job-search and skills-matching SaaS. It continuously harvests live job postings from employers' applicant-tracking systems, normalizes them (location, dates), extracts the skills each posting demands from an in-house skills ontology, and scores how well a candidate's résumé matches — with search, skill clouds, trending, and favorite/applied tracking on top. It is one of the constellation's paid web products: access and tiers are decided by GCStore Core, not by GCruiter.
Internal architecture
GCruiter is a Spring Boot backend plus an Angular single-page app over one relational schema, organised into cooperating subsystems:
- Capture & discovery. A connector per ATS platform (Greenhouse, Lever, Ashby, SmartRecruiters, Recruitee, Workable, Workday, Oracle Fusion) behind a common interface. A parametrized discovery sweeper probes company × platform to learn who-uses-what (easiest connector first), then mass-captures postings. Two lanes: a full sweep that also detects closures, and a delta fast lane that only ingests what's new. Long jobs are cancellable (a run registry + cooperative cancel flags) and commit in batches.
- Normalization. Geo resolution (free-text location → country / state / city, via an open gazetteer) and published-date parsing that heals timezone drift, both runnable as backfills over stored postings.
- Skills engine — the differentiator. A taxonomy ingested idempotently from open datasets — ESCO, O*NET, Wikidata, and Stack Overflow tags — matched by an in-memory gazetteer: a token n-gram index over the active aliases, with ambiguity/co-term rules so short collisions ("Go", "R", "C") only fire with a confirming term nearby. Extraction writes posting skills; the same matcher runs over résumés for matching.
- Résumé & matching. Upload/parse a résumé, compute its skills, plus per-résumé custom skills the user types that the taxonomy doesn't (yet) know. Match % is
|posting ∩ user| / |posting|, surfaced on search rows and the marks tabs.
- Search / query & per-user. Full-text search, a location-tree filter, skill clouds, company insights, trending, and generic per-user marks (favorites, applied) keyed by the authenticated user.
- Auth & admin. Two token surfaces (below), a web entry gate, and an owner-only admin console (company CRUD, capture/ingest triggers, and lexicon curation).
Design decisions worth calling out
- Inclusion-first skills hygiene. A term is matchable only when a trusted dictionary confirms it — the curated seed, the tech-native Stack Overflow tags, or an authoritative ESCO/O*NET domain classification; everything else stays dormant (kept, never deleted) until confirmed. Exclusion rules can't be exhaustive and degrade into a patchwork, so the only subtractive rule is a thin margin driven by a public English frequency list as a collision detector. The result is a clean lexicon with no hand-maintained blocklists.
- Curation as an independent layer. An admin can enable/disable/create/delete terms; a hand toggle pins the alias so the automatic hygiene never overrides it. Automation and curation coexist without clobbering each other.
- One matcher everywhere. "Does the taxonomy know this term?" is answered by the same gazetteer whether the term is typed by a user or read from a résumé — one code path, one answer, no drift between surfaces.
- Least-privilege data access. A single JPA datasource routes write transactions to an elevated account and everything else to a read-only account (a lazy-connection proxy sets the route after the transaction's read-only flag is known). Public read paths physically cannot mutate.
- Write-minimization at capture. Unchanged postings are never re-written; last-seen is stamped in bulk conditionally, so same-day re-runs write near-zero rows.
Interfaces
- ⇆ Keycloak (customers realm). OIDC Authorization Code + PKCE login from the SPA (one public client). The backend verifies Keycloak tokens locally against the realm key set to gate per-user data and the owner-only admin surface.
- ⇆ GCStore Core. On entry the SPA exchanges the Keycloak identity at the delivery endpoint for a short-lived, scoped app token (or a locked response carrying a paywall offer — the web entry gate). GCruiter's public API verifies that token offline (audience, expiry, key set) — no per-request callback. Entitlement tiers arrive as scopes (also projected by GCStore Core into Keycloak client roles); the app maps scope suffixes to its own feature vocabulary.
- ⇆ External public data. Read-only ingestion: employer ATS platforms (postings), ESCO / O*NET / Wikidata / Stack Overflow (skills), and an open geo gazetteer (locations).
- Shared data plane. Its own schema on the shared MariaDB, reached over the internal container network.
Tech stack
Backend: Java 21, Spring Boot, Spring Data JPA, Liquibase (versioned migrations), MariaDB; JDK-native Ed25519 verification (no third-party JWT library); jsoup for HTML-scraping connectors. Frontend: Angular (standalone components), oidc-client-ts for the OIDC/PKCE flow, TypeScript. Deploy: Docker Compose behind the host nginx that owns TLS; loopback-only service port; secrets in an env file with fail-fast guards — the constellation's platform conventions.
Status & roadmap
Production. The full loop runs end to end — capture → normalize → extract skills → search / match → mark — with the web entry gate (entitlement paywall on the delivery grant) and the admin lexicon curation live.
- Server-side per-feature enforcement across the user and admin surfaces lands as GCStore Core's entitlement projection matures (scope-driven gating runs on the client today).
- App-token verification moves from a config-pinned key to dynamic key-set fetch for seamless rotation.
- The skills ontology keeps widening; connector breadth keeps growing (iCIMS, Taleo next). The paywall's checkout flow is being wired on GCStore Core.