Pushing frontier models
We put the newest large language models under deliberate load and look for the points where they break. With benchmarks, error analysis and calibration we measure what a model can actually do.
Research at AI Værk
We also compete against strong teams at hackathons. That is how we know from our own work what a model delivers and where its limits are, before it runs in your business. We publish methods, data and results openly.
Fields
We put the newest large language models under deliberate load and look for the points where they break. With benchmarks, error analysis and calibration we measure what a model can actually do.
We train our own models and adapt large language models to specialist domains. Having trained a model yourself, you see where its strengths and weaknesses come from and choose the right one for each task.
At hackathons we build under time pressure what current models make possible. It keeps us fast and shows early which ideas hold up in practice.
Publications
Publication 01 · Prediction markets
A heuristic classification and comparative analysis of trader types on a prediction market.
The study classifies more than one million pseudonymous wallets from their observable trading behaviour. It shows how sharply market weight, profit, and forecast quality diverge across Retail, Informed Traders, Whales, Algorithmic Traders, and Market Makers.
Read the studyWe test it with you on your data, build a benchmark together or train a model of your own. We are just as glad to talk about research collaborations.