Every bank call checked, not just samples
The problem today
Every phone call between customers and bank advisors is recorded, but only a small part is ever reviewed. Key staff listen to samples. With thousands of call minutes per day, fraudulent behaviour such as insider trading mostly goes undetected, on both sides of the line. The calls are in Swiss German, which is where off-the-shelf speech tools stop working.
The expected outcome
Build a system that automatically checks every phone call in Swiss German, turns it into a reasoned compliance alert and knows when it is unsure. The prototype transcribes the call, runs several kinds of AI-based fraud checks, flags the suspicious passages with a reason, and has an adjustable threshold to trade off false alarms against missed cases.
Who it serves
Compliance officers, who review the suspicious passages directly instead of sampling. Banks, which meet their regulatory duties without gaps. And customers, who are protected against fraud. Inventx builds and runs the platforms these teams work in; Outcept brings the case and is on site all weekend.
Data
- Test calls: audio, WAV, mono, 16 kHz, in Swiss German. 60 synthetic calls from 30 dialogues, each in a clean and a noisy version, about 4 hours 40 minutes. Plus 8 calls recorded by two people. - Reference transcripts: TXT, the dialogue for every call, turn by turn. They are the scripts the recordings were made from, so no team gets stuck on transcription. For the recorded calls the spoken wording can differ from the script. - Expected assessments: per call alarm, no_alert or review, with the reasoning and the passages that carry the decision. - Keyword list: JSON, preliminary keyword families such as insider trading, access information or disclosure to third parties. - Hidden test set: about 70 percent of the material goes to the teams, Inventx and Outcept keep 30 percent back. On Sunday the solutions run on that part, so precision and recall are fairly comparable. - Delivery: a ZIP with a README, handed out at the kickoff. Link to the files: GitHub All calls are fictional. There is no real customer or bank data. Keyword list and expected assessments are test conventions for the sprint, not official rules of Inventx.
Technology
The audio comes from a sensitive banking context, so the solution must be able to run in a secured environment. Rule: if it can be self-hosted, it is allowed. - Allowed: Whisper or Swiss German variants, self-hosted. Open language models such as Llama or Mistral. Any model you can run on the GPUs you get. Any programming language and any framework. - Not allowed for the core pipeline: pure cloud services without a self-hostable model, such as Deepgram, AssemblyAI or ElevenLabs. - Compute: you get GPUs for the weekend. Note on the event-wide credits: the OpenAI and ElevenLabs credits available to every builder do not fit this rule. Do not build transcription or the fraud checks on them.
Core use case
In the target picture the checks are integrated into the bank's call recording solution and run automatically on every call, without anyone uploading anything. A call comes in as audio. The system transcribes it, runs several kinds of AI-based fraud checks (for example for insider trading, for sharing access information or for disclosure to third parties) and decides: alert, review or no alert. If a call exceeds the threshold, the responsible compliance officer receives the suspicious audio snippets with a timestamp and a reason; the full call is optional. They want to hear the suspicious passage, not a 20-minute call. Where the system is unsure, it says so and hands the case to a human instead of guessing.
Business value
Full coverage instead of samples, uniform criteria instead of individual judgement, far less manual review time, and traceable decisions for audits and the regulator.
Optional stretch goals
- Speaker separation: tell customer and advisor apart. - Notifying the right person on an alert. - Feedback: mark false alarms, and the system suggests better thresholds. - Patterns across several calls, such as an advisor who stands out repeatedly. - An operating concept for running this in production in a secured environment. - The Outcept trigger API: a small local test API your solution reports its hits to. POST sends a trigger (url, title, optional description), GET lists all triggers. Ideally the url links back into your product, straight to the suspicious passage with about 10 seconds before and after. Code and docs: github.com/Outcept/trigger-api None of these are required for a valid submission.
Presentation format
Team pitches on Sunday, 09:00 to 11:00: 5 minutes live demo, 3 minutes questions. Slides are optional. Final submission on Sunday at 08:30.
Key presentation elements
One call from audio to alert, with the suspicious audio snippets. One change to the threshold or the keyword list with a visible effect on the number of hits. One borderline case: a harmless call that still contains a keyword. Your detection quality on the test set.
Prototype requirements
Required for a valid submission: 1. Transcription of the Swiss German call. 2. Several kinds of AI-based fraud checks. 3. Suspicious passages as audio snippets, each with a timestamp and a reason. 4. An adjustable threshold. 5. A metric for false alarms and missed cases, reported on the test set. The prototype runs without manual steps per call. Everything else is a stretch goal.
How the jury scores
| Criterion | Weight | What we reward |
|---|---|---|
| Creativity & innovation | 20% | A fresh, surprising and workable approach. |
| Feasibility | 10% | Realistic architecture that could run in a bank's secured environment, on self-hostable models. |
| Design & usability | 25% | Clear and simple to use. A compliance officer understands each alert and its reason, and can adjust threshold and keyword list without code. |
| Impact & relevance | 20% | Does it solve the actual problem effectively? |
| Detection quality & traceability | 25% | Few false alarms and few missed cases on the hidden test set (precision and recall). Every hit comes with its passage and criterion. The solution shows when it is unsure and escalates instead of guessing. |
Prize
Goodies from Inventx and Raspberry Pis from Outcept for the winning team of the case.