Breakthrough with Data or AI – International – Bank of Georgia – DS Agent

Bank of Georgia’s DS-Agent is changing the economics of data science delivery. In its first benchmarked project, two people and an AI agent crew completed work estimated to require four people and 12 weeks in just eight weeks, delivering a threefold efficiency gain while producing models judged better than the conventional approach. 

AI copilots can make data scientists faster, so Bank of Georgia asked the fundamental question: what if AI could change the composition of the data science team itself? 

The result is DS-Agent, an internally developed multi-agent system in which humans provide direction, experience, and judgement, while specialised AI agents undertake much of the planning, investigation, development, validation, and deployment. 

 

From Copilot to Agent Crew 

DS-Agent emerged from a capacity problem. Demand for Bank of Georgia’s data analytics team consistently exceeded what it could deliver, but simply adding people risked increasing coordination overhead alongside capacity. Rather than make individual data scientists incrementally more productive, the team redesigned the delivery model around small human-led pods supported by an AI crew. 

More than 10 specialised agents and micro-agents perform different roles. An Orchestrator directs activity, while other agents research context, undertake modelling, validate findings, challenge reasoning, enforce quality gates, translate business requirements, and perform supporting tasks. 

Humans remain responsible for direction and decisions and, crucially, the architecture prevents the system from passing important milestones without explicit human approval. The philosophy is “humans lead, agents execute.” 

 

“Great! Data scientists creating agents to help data scientists (and the business)…Very good example of specialists drinking their own champagne with benefits.” – Judges’ comments 

 

Multiplying Capacity, Not Just Speed 

The first benchmarked project demonstrated the potential. A sales propensity modelling project estimated to require four people over 12 weeks was completed by two people working with DS-Agent in eight weeks. 

The results included: 

  • Half the human project team, falling from four people to two. 
  • Delivery in eight weeks rather than the estimated 12. 
  • Approximately threefold improvement in project efficiency when time and headcount are combined. 
  • Faster time to production. 
  • Final propensity models judged better than the conventional baseline. 
  • Two projects now delivered through DS-Agent, with six more in development. 

Speed has changed how the team experiments. Hypotheses that would previously consume weeks can be investigated in hours, allowing data scientists to explore approaches that might otherwise have been rejected as too costly or time-consuming. 

As senior data scientist Nikoloz Jangisherashvili explains: “Before DS-Agent, a failed experiment was a wasted week. Now it’s an afternoon I’d have spent in meetings anyway.” 

 

Engineering for AI’s Failure Modes 

Making agentic AI work on real analytical projects required more than better prompting. Bank of Georgia built an internal LLM Wiki that converts fragmented organisational knowledge into a searchable knowledge graph, giving agents access to relevant institutional context. 

Automated safeguards check work before and after significant actions. An external memory system prevents agents losing important findings during long-running projects. Human review is deliberately introduced as friction where judgement matters, preventing employees from accepting AI outputs too readily. 

A self-hosted board server provides shared memory and coordination across the agent crew, while giving data scientists a dashboard through which they can review findings, add comments, approve decisions, and remain involved without interrupting autonomous workflows. 

Each element emerged in response to failure modes encountered during real project execution which is what makes DS-Agent more significant than task automation. 

The system is attempting to automate elements of the investigate-challenge-verify-synthesise cycle at the heart of analytical work, while preserving human accountability for the decisions that matter. 

The breakthrough is demonstrating a different unit of analytical delivery where a small human team directing a scalable crew of specialised AI agents have greater capacity for experimentation, deeper analysis, and human judgement built into the architecture. 

DataIQ Awards 2026
Year: 2026
Category: Breakthrough with Data or AI - International

Upcoming DataIQ events

No event found!