CRASH

Cultural Heritage Advanced Search Hackathon

WHAT
two days of hacking on search in cultural heritage data
WHEN
30 November and 1 December 2026
WHERE
Graz, Austria
COST
free, no participation fee
MOTTO
Vibes and Checks
register --by 1-november

Why

Search is at the heart of research. A researcher knows what a good search tool does for them: it lets them move through a collection, it lets them steer the result, it is fast, and it works on the next dataset as well. Research software often falls short at this point.

Search is retrieval, and what a tool retrieves reflects the question: the results are the items most similar to the query. Similarity is the first decision in every search tool, and most tools make it for the user and hide it. But finding is only the start. A good tool also shows why a result is there. It lets the user change what similar means and see the results follow. It works for people who read with a screen reader or in another language. It says what one search costs in time and energy. And now and then it finds the thing nobody asked for. CRASH is about tools that do this work in the open, so that the user can see it and steer it.

For two days, CRASH puts cultural heritage data and the people who care about it in one room.

Read the manifesto

Search is the point

How we search, and how we handle what we find, shapes how we think and work. A search interface brings many decisions together in one place: how the data is processed, how it is shown, how a person can take it in, and who has a say in all of this.

The largest of these decisions is what counts as similar. Every search ranks its results by some measure of closeness between the question and the items. The same measure decides what a user sees next to a result, what a tool recommends, and what it leaves out. In the digital humanities, this measure is a scholarly choice. Two charters, two images or two texts can be similar in many ways, and each way answers a different question. Sometimes the useful result is the dissimilar one, the item that breaks the pattern. How to model similarity has been a question of information retrieval since its beginnings, and it is still one of its most persistent ones. A search interface that lets the user see and change this choice is rare.

Ranking is one part of a search tool. Whether a person can trust a result, steer the tool, use it at all, and afford to run it, are the others. Research software and research infrastructures have their gaps in all of these places.

What changed

Large language models and agents have reached many professions, and cultural heritage is one of them. Everybody can build a search application in an afternoon now. We think this is good, and we will build in the same way at this hackathon, with code assistants and agents, used with care. We want actual hacking: work towards a target, and sometimes work for the pleasure of building. Culture is fun, and the digital humanities are the academic way to have that fun. We are interested in the specific and the weird: the objects that are new even to the experts who work with them every day.

What goes wrong without shared ground

Everybody can build their own tool, and many do. But the oldest question of software remains: who is this for, and why? The people who use these tools know their limits by now. Answers can be invented. The way to a result is hidden. The computation is expensive. The same question gets a different answer on every run. The models show their reasoning, but that reasoning explains little; it mostly reassures.

If every researcher works with their own agents, every agent follows its own rules. Without shared rules, everybody arrives at their own truth, and we lose the ground on which we can compare and argue.

What a check is

Cultural heritage has strong quantitative methods by now. They give us numbers for what counts as similar or as relevant. But relevance is a matter of definition. It depends on the signal we measure and on the context of that signal. Research on recommender systems has developed many measures beyond accuracy: diversity, novelty, coverage, fairness. The same is true for the other qualities of a tool. Whether an explanation helps, whether a control is understood, whether a screen reader gets through, and what a query costs, can all be tested.

So everything built at CRASH answers to a check. A check is a test that a team designs for its own tool: what does it claim, and how can that claim be shown? We check ourselves too. We measure the ecological footprint of the event, in CO2 and in compute, and we publish the numbers.

What we hope comes out

Working prototypes, on real material, with a test beside each of them. A habit of saying how a result was checked. And people who know each other afterwards and keep building. On Tuesday the room and the jury each rank the prototypes, and the first three places in each ranking get a prize. Everybody leaves with the same room, the same food and the same clusters behind them. Bring your data, part of your team, and good vibes. We bring the checks.

Vibes and Checks

Vibes

Build what you want, with your team, your data and your problem.

Winner of hearts, chosen by the room.

Checks

Pick one thing to do well, and show at the end how you checked it.

Winner of numbers, chosen by the jury.

Every team picks one focus area. In the final presentation, the team shows one check for that area: what it tested, and how. There are two rankings, one by the room and one by the jury, each with places one to three. The prizes are small and symbolic.

AreaQuestionExample of a check
EfficiencyWhat does one search cost?time, compute and energy per query
TrustCan I see why this result is here?every result links to its record; the same query gives the same answer twice
ControlCan the user steer the outcome?the user changes what similar means, and the results follow
AccessWho can use it?screen reader, voice, plain language, a second language
DiscoveryDoes it find the weird things?measures beyond accuracy: diversity, serendipity, dissimilar items

Agents are welcome in every area. How you build is your choice.

Topics

Some examples of methods we would like to see:

  • retrieval over image, text and layout at the same time
  • embeddings for every kind of object, and an account of what they miss
  • re-ranking that improves a result list without burying it
  • retrieval-augmented generation (RAG) that cites its sources and says where it found nothing
  • agents that search on a user's behalf and keep a log of what they did
  • vision-language models that read handwriting
  • search across languages, scripts and the eight spellings of one village
  • detection of copies, forgeries and near-duplicates
  • recommendation and exploration as an alternative to the query box
  • knowledge graphs and linked open data
  • evaluation beyond accuracy
  • the cost of one query in watts

Some examples of material:

  • charters with and without seals
  • manuscripts with their marginalia and doodles
  • newspapers with text recognition from 2009
  • letters, diaries and account books
  • maps that disagree with each other
  • museum objects described in three words
  • photographs without captions
  • audio, film and 3D scans
  • catalogue cards that describe other catalogue cards
  • metadata in every standard, and in none

Who can come

We invite students, researchers, practitioners and experts from academia, memory institutions and industry, for example:

  • hackers
  • people from natural language processing (NLP) and computer vision (CV)
  • data stewards
  • curators
  • archivists
  • librarians
  • designers

The list is not complete. If the event interests you, register.

The most valuable thing you can bring is your team, or part of your lab, with a dataset and a problem you know well. If that is not possible, come on your own. You then bring something the room can work with: a skill, a dataset, or a problem. You do not have to code. We especially encourage people of genders underrepresented in tech to register.

The invitation goes to the institutions of CLARIAH-AT, to partner universities, and to CLARIN and DARIAH institutions in the countries around Austria and beyond. Places are limited. We read every registration, and we put the room together so that the teams are strong and mixed.

The event language is English. A little German in the team helps, because some datasets have German in them. Hacking is on site only.

Teams

Come prepared, and bring part of your lab:

  • one or two coders
  • one or two people from the domain, from human-computer interaction (HCI) or from design
  • a dataset or a problem you know well

A team that arrives with its question can spend the whole two days on hacking. A team has two to five people. Everybody registers alone; a team is accepted as a whole and comes as it has registered.

If you have no team yet, register anyway and say that you are looking for one. You then get the link to a chat room where people find partners and teams. When you have a team, add its name to your registration; you can change your registration until 1 November. On Monday, a pitch can still end with "we need", and you can join a team there.

The teams work on different problems, but they share one room. Asking the team next to you, showing them what you have, and passing on a trick are part of the two days. The winners get small prizes and the attention of the room. Everybody else gets the same room, the same food and the same clusters.

What we bring

  • rooms, and 15 or more hours of hacking time, depending on your endurance;
  • the DHinfra clusters, run in Graz, with current open-weights models hosted for you and the graphics processors (GPUs) of the clusters at your disposal. Bring your own laptop;
  • two impulse talks, one by a domain expert and one by a technical expert;
  • warm and cold food, snacks, coffee and tea throughout both days;
  • one special and a few more general prepared datasets on Austrian cultural heritage, announced by 15 November;
  • support for travel and accommodation, for those who need it.

Your dataset

Prepare your own data and your own problem, and describe the dataset in the registration. It does not have to be large or famous. It needs modalities we can work with, and it should be ready or nearly ready. A link to the dataset, or to a large subset of it, helps us confirm that it can be worked with.

If you can share the dataset with the other teams during the event, and if others may reuse it afterwards, that is a bonus. Teams with a dataset come first when places are short.

If you bring no dataset, you can work on one of ours. We announce them by 15 November, so prepare your own problem in the meantime.

What the registration asks about your dataset (if you have one)

Four questions, answered once per team by its contact person.

Nameof the dataset
What it is, and what you fail to find in it todaya few sentences. Where you know it, mention material, period, institution, size, language and modalities
Link or pointerto the dataset or to a large subset, if you have one
Rightstick what applies: the other teams may use it during the event; others may reuse it afterwards; we may list it on this page; results may be published

Schedule

Preliminary. The start and the end of both days are fixed; the times in between may still move.

Sunday 29 November

Arrival, and a get-together from 18:00 at your own cost. Registered participants get the place by mail.

Day 1, Monday 30 November

09:00Registration
09:30Welcome
10:00Impulse talk by a domain expert: what do cultural heritage experts hope AI can do for searching and finding material?
11:15Pitches, two minutes per team
12:00Hacking starts
20:00Impulse talk by a technical expert. The venue stays open into the night

Day 2, Tuesday 1 December

08:00Doors open
15:00End of hacking
15:30Presentations, three minutes per team
16:30Award ceremony
17:00End

The two impulse talks are streamed live with their Q&A. Nothing is recorded.

Dates

1 Novemberregistration closes
4 Novemberconfirmations, with your support for travel and accommodation
15 Novemberdatasets published
23 Novembercluster accounts and access
29 Novemberarrival, get-together
30 Novemberday 1, from 09:00
1 Decemberday 2, until 17:00

All times are Central European Time (CET). A deadline ends at 23:59 on its day.

Travel and accommodation

Taking part is free of charge: there is no participation fee. You only need to make sure that you can travel to Graz and stay for both days.

Everybody can ask for support for travel and accommodation, up to 150 EUR per person. Say in the registration where you travel from, what it costs, and whether you can come without support. You get your amount with the confirmation on 4 November, so that you can book. The money is paid after the event.

Accommodation covers the night from Monday to Tuesday for everybody who comes from outside Graz. It covers the night from Sunday to Monday if you cannot reach Graz by 09:00 on Monday otherwise. It covers the night from Tuesday to Wednesday if you cannot travel home after 17:00 on Tuesday.

Conduct

Be kind to each other. Everybody in the room cares about this material, and that is enough common ground.

The code of conduct

It holds in the rooms, at the evening, in the chat rooms and in the code.

  • Come prepared, and help the people around you get up to speed.
  • Questions are more than welcome.
  • No harassment, no discrimination, no unwanted attention.
  • Ask before you photograph or record somebody.
  • Credit what you use. Data comes with licences and with people; keep both intact.
  • Try to say what a model did and what you did.
  • Look after each other, also in the evening.

If something is wrong, tell an organiser on site or write to crash.organizers@gmail.com. We treat it as confidential and decide the next step together with you. The organisers can ask a person to leave the event.

Photos

We take photos of the event. In the registration you say whether photos may show you. We do not publish a photo of anybody who said no.

Footprint

We measure the footprint of the event, in CO2 and in compute, and we publish it.

Venue

UNICORN Startup & Innovation Hub
Schubertstraße 6a, 8010 Graz, Austria
Top floor, conference deck.

The building is barrier-free, with a lift and automatic door openers. A second place for hacking is close by: the meeting and office rooms at Elisabethstraße 59/III.

Register

open until 1 November 2026

open the registration form (opens in a new tab)

Register early. When places or travel money are short, earlier registrations and teams with a dataset come first. Everybody else goes on a waiting list.

After the event, we show the results here: the teams and what they built. We ask every team to share its code under Apache-2.0 or MIT, in its own repository or under crash-this. Copyright stays with the authors.

Organisers

Florian Atzenhofer-Baumgartner, Chiara Citro, David Fleischhacker, Bernhard Geiger, Hussain Hussain, Roman Kern, Dominik Kowald, Elisabeth Raunig and Thea Schaaf.

We work at the Department of Digital Humanities of the University of Graz, at Graz University of Technology and at Know Center.

CRASH is funded by CLARIAH-AT. For questions, and for sponsoring, write to crash.organizers@gmail.com.