The challenge
- No GPU estate, no container platform and no integration layer to core banking.
- A board expectation of production within the financial year, without a data-residency exception.
- Data residency: SBP rules and the bank's own policy prohibited customer data leaving controlled infrastructure. The pilot's hosted model APIs could not be used in production.
- GPU capacity: inference and fine-tuning needed pools sized, scheduled and isolated correctly without breaking the cost case or the customer channel.
- Legacy integration: core banking, CRM and the contact-centre platform had never exposed APIs to an agent layer. The integration surface had to be built and secured.
- Auditability: internal audit and SBP expected every model, deployment and access event to be logged, attributable and reproducible. The pilot had none of this.
Requirements
Agreed with the bank's architecture board, CISO and internal audit before design began.
Data residency
All customer data, prompts, embeddings and logs remain in bank-controlled infrastructure in Pakistan. No external model APIs.
Availability
99.9% monthly availability for customer-facing services; RTO 4 hours, RPO 15 minutes.
Auditability
Immutable logs for every platform, access and model event, exportable to the bank's SIEM and to regulators on request.
Capacity
GPU capacity for production inference and periodic fine-tuning, with utilisation visible per workload.
Integration
Secure API access to core banking, CRM and the contact-centre platform through a governed gateway.
Ownership
Documentation, training and a staged handover so the bank's team can operate the platform without dependency.
Our approach
Cloudlit engineers worked on site with the bank's infrastructure and security teams through five phases.
- Assess
Current-state review of infrastructure, network, identity and compliance posture. Gap register against target architecture.
- Design
Reference architecture, sizing, security controls and runbooks agreed with the bank's architecture board and CISO.
- Build
Landing zone, OpenShift platform, CI/CD, monitoring and DR provisioned as code in the bank's environment.
- Operate
Cloudlit runs the platform to agreed SLAs with the bank's team alongside: on-call, patching, capacity, incident response.
- Transfer
Documentation, training and staged handover so the bank's own team can own the platform without dependency.
How the delivery ran
Eight stages from first meeting to steady-state operation. The bank signed off at three gates (Approval, Test and Go-live) before anything reached production.
- Discovery: requirements, constraints, stakeholders. Deliverable: scope document.
- Assessment: current-state infrastructure and security review. Deliverable: gap register.
- Design: target architecture, sizing, controls. Deliverable: architecture pack.
- Approval (bank gate): architecture board and CISO sign-off. Deliverable: signed design.
- Build: provision as code, pipelines, monitoring. Deliverable: working environment.
- Test (bank gate): security testing, DR failover, UAT. Deliverable: test evidence.
- Go-live (bank gate): change approval, cutover, hypercare. Deliverable: production release.
- Operate: SLA-backed operations and reporting. Deliverable: monthly service report.
Why it worked
Requirements before architecture
The CISO and internal audit defined the control set first. Architecture was designed to satisfy it, so nothing needed retrofitting at the approval gate.
Everything as code
Infrastructure, platform and pipelines were declared in version control from day one. Every change was reviewable, repeatable and reversible.
Bank engineers in the room
Cloudlit worked on site with the bank's team rather than delivering a finished box. Handover was a phase of the project, not an afterthought.
Solution architecture
Five layers, one control plane. Customer data, models, prompts and logs never leave the bank-controlled perimeter; only the placement of each layer changes between deployment models. Layers 02 to 05 run on bank infrastructure.
- Channels
WhatsApp, mobile app, web, contact centre, branch. Customer and staff entry points with no direct path to bank systems.
- Access
API gateway, WAF, DDoS protection, identity broker and rate limits. Every request authenticated, authorised and logged before it enters the platform.
- Application
Agent orchestration, business logic, session state and policy checks. A governed application layer on OpenShift that decides what a model may see and do.
- AI platform
Model serving on GPU node pools, fine-tuning jobs, model registry and vector store. Containerised models under resource quotas, no external model APIs.
- Data and core systems
Core banking, CRM, contact-centre platform, encrypted stores and an immutable audit log. Bank-held keys (KMS/HSM) and integration through a governed internal gateway.
Cross-cutting controls on every layer
Identity
SSO, MFA, RBAC and PAM.
Security
Segmentation, signed images and admission policy.
Observability
Metrics, logs and traces to the bank SIEM.
GitOps
Everything declared in version control.
Resilience
DR site, tested failover, RPO 15 minutes.
Sovereignty
Data and keys stay with the bank.
Platform build
Red Hat OpenShift delivered as a production platform, not a cluster install. Eight engineering areas, each with a defined deliverable and the evidence the bank's auditors receive. The stack is provisioned bottom-up as code: hardware, RHEL CoreOS and OpenShift, platform services, then workloads.
Architecture
Cluster topology, node pools and multi-site layout: control plane on three nodes, dedicated GPU node pools, active/DR site design.
Applications
Containerisation and operator-based deployment: Helm charts per service, operators for stateful services, blue/green promotion.
Scalability
Autoscaling and GPU scheduling: HPA and cluster autoscaler, GPU quotas per workload, quarterly capacity model.
Storage and backup
Persistent storage and recovery: block and file storage classes, snapshot schedule, restore tested monthly.
CI/CD
GitOps delivery across environments: Tekton build and test, Argo CD sync to dev, UAT and prod, approval gates in Git.
Disaster recovery
Secondary site and failover: async replication with RPO 15 minutes, documented runbook, full failover test twice a year.
Monitoring
Health of platform and workloads: metrics, logs and traces, alert routing to the bank SOC, monthly service report.
Security
Controls built into the pipeline: image scanning and signing, admission and network policy, RBAC mapped to bank roles.
Infrastructure management
- Infrastructure as code: all infrastructure in version control, no console changes in production (Terraform, ARM/Bicep, CloudFormation).
- GitOps delivery: declarative sync through dev, UAT and production with approval gates (Tekton, Argo CD, OpenShift GitOps).
- Environment separation: isolated non-production; production data never copied down unmasked.
- Change and patch: scheduled windows, CVE response by severity, rollback tested before upgrades. Critical patches within 72 hours.
- Capacity: monthly review against growth; GPU utilisation and cost per inference per workload on dashboards.
Security and compliance
Six control areas implemented in the platform build and evidenced to the bank's CISO and internal audit.
Identity and privileged access
SSO with the bank's directory, MFA, role-based access, privileged session recording.
Encryption and key custody
Encryption at rest and in transit; keys held in the bank's KMS or HSM, never by Cloudlit.
Network segmentation
Zone-based network policies, east-west controls inside the cluster, egress allow-listing.
Workload security
Image scanning in the pipeline, signed images, admission control, runtime policy enforcement.
Logging and SIEM
Immutable audit logs for platform, access and model events, shipped to the bank's SIEM.
Vulnerability management
Continuous scanning, severity-based patch SLAs, independent penetration testing.
Operations
Service targets Cloudlit commits to for the bank's production environment.
Enterprise monitoring
24×7 monitoring of infrastructure, platform and application health. Critical patches within 72 hours of vendor release.
Disaster recovery
Secondary-site replication and documented failover. Full DR test twice a year, results reported to the bank.
Incident response
P1: 15-minute response, 4-hour resolution target. P2: 1-hour response. Post-incident review within 5 business days.
Reporting
Monthly service report covering availability, incidents, changes, capacity and patch status, in the bank's governance format.
Use case: connecting the bank's on-site environment with the cloud
Customer-facing AI services scale in a single-tenant private cloud tier. Models, customer data and core-banking integration stay in the bank's own data centres. The two sides are joined by dedicated, encrypted connectivity with automatic failover, and no public path exists between them.
- Ingress
Traffic terminates at the cloud edge: WAF, DDoS protection, API gateway. Nothing from the internet reaches the bank directly.
- Authenticate
The identity broker validates the customer or agent; the token is checked again at the bank's internal gateway.
- Transit
The request crosses the dedicated circuit; BGP fails over to IPsec VPN in seconds. Both paths are encrypted.
- Process
Models run on GPU node pools in the bank data centre and read core data through the governed gateway.
- Respond
Only the response returns to the cloud tier. No prompt, log or customer record is written outside the bank.
- Operate
One monitoring plane across both sides; logs to the bank SIEM; monthly service report.
Why the split is where it is
Front end in the cloud
Conversational load is bursty. Autoscaling stateless tiers avoids sizing on-prem for peak.
Models on-premise
Weights, prompts and embeddings are derived from customer data and carry the same residency obligation.
Keys with the bank
Cloud key vault and on-prem HSM both hold bank-controlled keys. Cloudlit never holds key material.
Two independent paths
Circuit plus VPN removes a single carrier as a point of failure without opening a public route.
Engagement team
The roles assigned to every banking engagement and what each is accountable for.
Project manager
Single point of accountability for scope, schedule, risk and reporting. Runs the sign-off gates with the bank's PMO: weekly status and RAID log, change and budget control, escalation owner.
Cloud engineers
Design and build the platform: infrastructure as code, container platform, CI/CD, monitoring and DR. Landing zone and network build, OpenShift and GPU node pools, pipelines and observability.
Security engineers
Own the control set: identity, encryption, segmentation, hardening, logging and audit evidence for the bank's CISO. IAM, PAM and key custody, vulnerability management, SIEM integration and audit packs.
How the team works with the bank
On site inside the bank's premises under the bank's access policies. Weekly steering with the bank's PMO; architecture board and CISO sign-off at each gate; monthly service report in the bank's format. Documentation, training and staged transfer so the bank's own team can operate the platform without dependency on Cloudlit.
What Cloudlit delivered
Dedicated GPU node pools on OpenShift
GPU capacity provisioned inside the bank data centre, scheduled through the container platform with quotas per workload and utilisation and cost visible monthly.
Governed internal gateway
Service mesh and message queue in front of core banking, CRM and the contact-centre platform. Every call authenticated and logged.
Immutable audit trail
Platform, access and model events written to immutable storage, exported to the bank SIEM and to regulators on request.
Landing zone, CI/CD, monitoring and DR as code
Provisioned in the bank's environment through Tekton and Argo CD with approval gates in Git and a tested secondary site.
SLA-backed operations and handover
Cloudlit operates the platform alongside the bank's infrastructure team under agreed SLAs, with documented handover in progress.
Results
| Problem | At pilot | In production |
|---|---|---|
| Data residency | Customer data sent to hosted model APIs outside bank control. Blocked by SBP rules and bank policy. | All customer data, prompts, embeddings and logs inside the bank perimeter. No external model APIs. No residency exception needed. |
| GPU capacity | No GPU estate. Inference depended on a vendor; no view of cost per inference. | Dedicated GPU node pools on OpenShift with quotas per workload; utilisation and cost visible monthly. |
| Legacy integration | Core banking, CRM and contact centre had never exposed APIs to an agent layer. | Governed internal gateway with service mesh and message queue; every call authenticated and logged. |
| Auditability | No attributable log of model, deployment or access events. | Immutable audit trail for every platform, access and model event, exported to the bank SIEM and regulators on request. |
| Ownership | Pilot operated by a vendor; bank team had no runbooks or access. | Bank infrastructure team operates alongside Cloudlit under SLAs, with documented handover in progress. |
- The platform is live on bank-controlled infrastructure with no data-residency exception.
- Service targets in the operating agreement are reported to the bank every month: 99.9% monthly availability, 4-hour RTO with a full DR test twice a year, 15-minute RPO with async replication to the DR site, 15-minute P1 response with a 4-hour resolution target, and critical patches within 72 hours of vendor release.
Scale across banking deployments
Infrastructure Cloudlit provisions and operates for AI platforms in banking. Application-level metrics belong to the client institutions.
- 20+
- Production clusters
- 50+
- GPU nodes provisioned
- 7
- Institutions
- 3
- Deployment models