Why
Search is at the heart of research. A researcher knows what a good search tool does for them: it lets them move through a collection, it lets them steer the result, it is fast, and it works on the next dataset as well. Research software often falls short at this point.
Search is retrieval, and what a tool retrieves reflects the question: the results are the items most similar to the query. Similarity is the first decision in every search tool, and most tools make it for the user and hide it. But finding is only the start. A good tool also shows why a result is there. It lets the user change what similar means and see the results follow. It works for people who read with a screen reader or in another language. It says what one search costs in time and energy. And now and then it finds the thing nobody asked for. CRASH is about tools that do this work in the open, so that the user can see it and steer it.
For two days, CRASH puts cultural heritage data and the people who care about it in one room.
Read the manifesto
Search is the point
How we search, and how we handle what we find, shapes how we think and work. A search interface brings many decisions together in one place: how the data is processed, how it is shown, how a person can take it in, and who has a say in all of this.
The largest of these decisions is what counts as similar. Every search ranks its results by some measure of closeness between the question and the items. The same measure decides what a user sees next to a result, what a tool recommends, and what it leaves out. In the digital humanities, this measure is a scholarly choice. Two charters, two images or two texts can be similar in many ways, and each way answers a different question. Sometimes the useful result is the dissimilar one, the item that breaks the pattern. How to model similarity has been a question of information retrieval since its beginnings, and it is still one of its most persistent ones. A search interface that lets the user see and change this choice is rare.
Ranking is one part of a search tool. Whether a person can trust a result, steer the tool, use it at all, and afford to run it, are the others. Research software and research infrastructures have their gaps in all of these places.
What changed
Large language models and agents have reached many professions, and cultural heritage is one of them. Everybody can build a search application in an afternoon now. We think this is good, and we will build in the same way at this hackathon, with code assistants and agents, used with care. We want actual hacking: work towards a target, and sometimes work for the pleasure of building. Culture is fun, and the digital humanities are the academic way to have that fun. We are interested in the specific and the weird: the objects that are new even to the experts who work with them every day.
What goes wrong without shared ground
Everybody can build their own tool, and many do. But the oldest question of software remains: who is this for, and why? The people who use these tools know their limits by now. Answers can be invented. The way to a result is hidden. The computation is expensive. The same question gets a different answer on every run. The models show their reasoning, but that reasoning explains little; it mostly reassures.
If every researcher works with their own agents, every agent follows its own rules. Without shared rules, everybody arrives at their own truth, and we lose the ground on which we can compare and argue.
What a check is
Cultural heritage has strong quantitative methods by now. They give us numbers for what counts as similar or as relevant. But relevance is a matter of definition. It depends on the signal we measure and on the context of that signal. Research on recommender systems has developed many measures beyond accuracy: diversity, novelty, coverage, fairness. The same is true for the other qualities of a tool. Whether an explanation helps, whether a control is understood, whether a screen reader gets through, and what a query costs, can all be tested.
So everything built at CRASH answers to a check. A check is a test that a team designs for its own tool: what does it claim, and how can that claim be shown? We check ourselves too. We measure the ecological footprint of the event, in CO2 and in compute, and we publish the numbers.
What we hope comes out
Working prototypes, on real material, with a test beside each of them. A habit of saying how a result was checked. And people who know each other afterwards and keep building. On Tuesday the room and the jury each rank the prototypes, and the first three places in each ranking get a prize. Everybody leaves with the same room, the same food and the same clusters behind them. Bring your data, part of your team, and good vibes. We bring the checks.
Vibes and Checks
Vibes
Build what you want, with your team, your data and your problem.
Winner of hearts, chosen by the room.
Checks
Pick one thing to do well, and show at the end how you checked it.
Winner of numbers, chosen by the jury.
Every team picks one focus area. In the final presentation, the team shows one check for that area: what it tested, and how. There are two rankings, one by the room and one by the jury, each with places one to three. The prizes are small and symbolic.
| Area | Question | Example of a check |
|---|---|---|
| Efficiency | What does one search cost? | time, compute and energy per query |
| Trust | Can I see why this result is here? | every result links to its record; the same query gives the same answer twice |
| Control | Can the user steer the outcome? | the user changes what similar means, and the results follow |
| Access | Who can use it? | screen reader, voice, plain language, a second language |
| Discovery | Does it find the weird things? | measures beyond accuracy: diversity, serendipity, dissimilar items |
Agents are welcome in every area. How you build is your choice.
Topics
Some examples of methods we would like to see:
- retrieval over image, text and layout at the same time
- embeddings for every kind of object, and an account of what they miss
- re-ranking that improves a result list without burying it
- retrieval-augmented generation (RAG) that cites its sources and says where it found nothing
- agents that search on a user's behalf and keep a log of what they did
- vision-language models that read handwriting
- search across languages, scripts and the eight spellings of one village
- detection of copies, forgeries and near-duplicates
- recommendation and exploration as an alternative to the query box
- knowledge graphs and linked open data
- evaluation beyond accuracy
- the cost of one query in watts
Some examples of material:
- charters with and without seals
- manuscripts with their marginalia and doodles
- newspapers with text recognition from 2009
- letters, diaries and account books
- maps that disagree with each other
- museum objects described in three words
- photographs without captions
- audio, film and 3D scans
- catalogue cards that describe other catalogue cards
- metadata in every standard, and in none
Who can come
We invite students, researchers, practitioners and experts from academia, memory institutions and industry, for example:
- hackers
- people from natural language processing (NLP) and computer vision (CV)
- data stewards
- curators
- archivists
- librarians
- designers
The list is not complete. If the event interests you, register.
The most valuable thing you can bring is your team, or part of your lab, with a dataset and a problem you know well. If that is not possible, come on your own. You then bring something the room can work with: a skill, a dataset, or a problem. You do not have to code. We especially encourage people of genders underrepresented in tech to register.
The invitation goes to the institutions of CLARIAH-AT, to partner universities, and to CLARIN and DARIAH institutions in the countries around Austria and beyond. Places are limited. We read every registration, and we put the room together so that the teams are strong and mixed.
The event language is English. A little German in the team helps, because some datasets have German in them. Hacking is on site only.
Teams
Come prepared, and bring part of your lab:
- one or two coders
- one or two people from the domain, from human-computer interaction (HCI) or from design
- a dataset or a problem you know well
A team that arrives with its question can spend the whole two days on hacking. A team has two to five people. Everybody registers alone; a team is accepted as a whole and comes as it has registered.
If you have no team yet, register anyway and say that you are looking for one. You then get the link to a chat room where people find partners and teams. When you have a team, add its name to your registration; you can change your registration until 1 November. On Monday, a pitch can still end with "we need", and you can join a team there.
The teams work on different problems, but they share one room. Asking the team next to you, showing them what you have, and passing on a trick are part of the two days. The winners get small prizes and the attention of the room. Everybody else gets the same room, the same food and the same clusters.
What we bring
- rooms, and 15 or more hours of hacking time, depending on your endurance;
- the DHinfra clusters, run in Graz, with current open-weights models hosted for you and the graphics processors (GPUs) of the clusters at your disposal. Bring your own laptop;
- two impulse talks, one by a domain expert and one by a technical expert;
- warm and cold food, snacks, coffee and tea throughout both days;
- one special and a few more general prepared datasets on Austrian cultural heritage, announced by 15 November;
- support for travel and accommodation, for those who need it.
Your dataset
Prepare your own data and your own problem, and describe the dataset in the registration. It does not have to be large or famous. It needs modalities we can work with, and it should be ready or nearly ready. A link to the dataset, or to a large subset of it, helps us confirm that it can be worked with.
If you can share the dataset with the other teams during the event, and if others may reuse it afterwards, that is a bonus. Teams with a dataset come first when places are short.
If you bring no dataset, you can work on one of ours. We announce them by 15 November, so prepare your own problem in the meantime.
What the registration asks about your dataset (if you have one)
Four questions, answered once per team by its contact person.
| Name | of the dataset |
| What it is, and what you fail to find in it today | a few sentences. Where you know it, mention material, period, institution, size, language and modalities |
| Link or pointer | to the dataset or to a large subset, if you have one |
| Rights | tick what applies: the other teams may use it during the event; others may reuse it afterwards; we may list it on this page; results may be published |
Schedule
Preliminary. The start and the end of both days are fixed; the times in between may still move.
Sunday 29 November
Arrival, and a get-together from 18:00 at your own cost. Registered participants get the place by mail.
Day 1, Monday 30 November
| 09:00 | Registration |
| 09:30 | Welcome |
| 10:00 | Impulse talk by a domain expert: what do cultural heritage experts hope AI can do for searching and finding material? |
| 11:15 | Pitches, two minutes per team |
| 12:00 | Hacking starts |
| 20:00 | Impulse talk by a technical expert. The venue stays open into the night |
Day 2, Tuesday 1 December
| 08:00 | Doors open |
| 15:00 | End of hacking |
| 15:30 | Presentations, three minutes per team |
| 16:30 | Award ceremony |
| 17:00 | End |
The two impulse talks are streamed live with their Q&A. Nothing is recorded.
Dates
| 1 November | registration closes |
| 4 November | confirmations, with your support for travel and accommodation |
| 15 November | datasets published |
| 23 November | cluster accounts and access |
| 29 November | arrival, get-together |
| 30 November | day 1, from 09:00 |
| 1 December | day 2, until 17:00 |
All times are Central European Time (CET). A deadline ends at 23:59 on its day.
Travel and accommodation
Taking part is free of charge: there is no participation fee. You only need to make sure that you can travel to Graz and stay for both days.
Everybody can ask for support for travel and accommodation, up to 150 EUR per person. Say in the registration where you travel from, what it costs, and whether you can come without support. You get your amount with the confirmation on 4 November, so that you can book. The money is paid after the event.
Accommodation covers the night from Monday to Tuesday for everybody who comes from outside Graz. It covers the night from Sunday to Monday if you cannot reach Graz by 09:00 on Monday otherwise. It covers the night from Tuesday to Wednesday if you cannot travel home after 17:00 on Tuesday.
Conduct
Be kind to each other. Everybody in the room cares about this material, and that is enough common ground.
The code of conduct
It holds in the rooms, at the evening, in the chat rooms and in the code.
- Come prepared, and help the people around you get up to speed.
- Questions are more than welcome.
- No harassment, no discrimination, no unwanted attention.
- Ask before you photograph or record somebody.
- Credit what you use. Data comes with licences and with people; keep both intact.
- Try to say what a model did and what you did.
- Look after each other, also in the evening.
If something is wrong, tell an organiser on site or write to crash.organizers@gmail.com. We treat it as confidential and decide the next step together with you. The organisers can ask a person to leave the event.
Photos
We take photos of the event. In the registration you say whether photos may show you. We do not publish a photo of anybody who said no.
Footprint
We measure the footprint of the event, in CO2 and in compute, and we publish it.
Venue
UNICORN Startup & Innovation Hub
Schubertstraße 6a, 8010 Graz, Austria
Top floor, conference deck.
The building is barrier-free, with a lift and automatic door openers. A second place for hacking is close by: the meeting and office rooms at Elisabethstraße 59/III.
Register
open until 1 November 2026
open the registration form (opens in a new tab)
Register early. When places or travel money are short, earlier registrations and teams with a dataset come first. Everybody else goes on a waiting list.
After the event, we show the results here: the teams and what they built. We ask every team to share its code under Apache-2.0 or MIT, in its own repository or under crash-this. Copyright stays with the authors.
Organisers
Florian Atzenhofer-Baumgartner, Chiara Citro, David Fleischhacker, Bernhard Geiger, Hussain Hussain, Roman Kern, Dominik Kowald, Elisabeth Raunig and Thea Schaaf.
We work at the Department of Digital Humanities of the University of Graz, at Graz University of Technology and at Know Center.
CRASH is funded by CLARIAH-AT. For questions, and for sponsoring, write to crash.organizers@gmail.com.