SkinVision tests Polteq iQ®: automated, data-driven robustness testing for complex AI systems
How do you test the robustness of an AI system?
SkinVision is a leading provider of early skin cancer detection through mobile devices. With more than 1.2 million users in over 50 countries worldwide, SkinVision is helping fight skin cancer on a global scale. Its mobile app can detect signs of melanoma, squamous cell carcinoma, basal cell carcinoma, and precancerous lesions. To date, SkinVision has helped identify more than 40,000 cases of skin cancer worldwide.
SkinVision conducted a pilot with Polteq iQ® to investigate how changes in lighting conditions affect the reliability of AI assessments.
An AI system can be impressively accurate. But what happens when conditions change? Does the system continue to deliver the same level of quality?
For SkinVision, this is not a theoretical question. The SkinVision mobile application helps users around the world detect potential skin cancer at an early stage. The app analyzes images of skin lesions using a deep learning algorithm and provides recommendations based on the results.
With more than 1.2 million users in over 50 countries, reliability is essential. The app must perform well not only under ideal conditions, but also across different smartphones, skin types, and lighting conditions.
That is why SkinVision wanted to investigate what happens to the accuracy of its algorithm when the color temperature of the lighting changes.
From traditional testing to large volumes of data
With traditional software, you can often determine in advance which input to provide and what result to expect. With a deep learning model, things work differently. The system is probabilistic and can produce different results for similar inputs.
This makes it difficult to demonstrate the robustness of an AI system using traditional testing methods alone.
SkinVision therefore engaged Polteq for a pilot using Polteq iQ®, Polteq’s proprietary framework for structured and automated testing of AI and ML systems.
The goal was not to test just one or a few scenarios, but to use large volumes of data to identify patterns and differences.
From question to insight in three days
During a three-day assessment, Polteq and SkinVision worked through several steps:
- analyzing the business requirements;
- developing an appropriate test strategy;
- setting up a test environment;
- connecting Polteq iQ® to an isolated version of the SkinVision API;
- executing large numbers of tests using synthetic data;
- collecting and analyzing the results;
- reporting and discussing the findings with the SkinVision team.
By automating the tests, a much larger volume of data could be processed than would have been possible through manual testing.
Not just: does it work? But also: does it keep working?
That is an important distinction when testing AI.
With a traditional system, regression testing often focuses on whether a change has affected existing functionality. With AI systems, you also want to know whether the model’s behavior changes and, more importantly, what that change means for quality.
Polteq iQ® makes it possible to compare different versions of an AI system. This makes the impact of a change on the results visible before a new version goes live.
In the SkinVision pilot, this approach was used to investigate the robustness of the assessment algorithm under variations in color temperature.
Testing AI requires a different way of thinking
The pilot confirmed for both SkinVision and Polteq that testing complex AI systems requires more than traditional test cases and expected outcomes.
When an AI model learns from data, the risks also change. These include variations in input, unexpected outcomes, bias, and changes in model behavior. To gain control over these risks, structured, automated, and data-driven testing methods are needed.
The collaboration with SkinVision demonstrates what such an approach looks like in practice: not simply assessing whether an AI system works well today, but investigating whether it remains reliable when conditions change.
From developing AI to controlling AI
AI is making software development faster. This makes the need for demonstrable quality more important than ever.
With Polteq iQ®, Polteq helps organizations make changes in AI systems measurable and understand their impact. This provides greater control over quality, particularly as the technology itself becomes increasingly unpredictable.
Want to learn more about Polteq iQ® and testing AI systems? Get in touch with us.