Showing posts with label APIs. Show all posts
Showing posts with label APIs. Show all posts

Sunday, 13 November 2011

On releasing museum data and the importance of licenses

I've been preparing for the workshop on 'Hacking and mash-ups for beginners' I'm running at the Museum Computer Network conference (MCN2011) this year, which as always means poking around the GLAM APIs, linked and open data services page for some nice datasets to use in exercises.  Meanwhile, people have been using NMSI data at Culture Hack North this weekend, and a question from that event made me realised I never blogged here about the collections data released by NMSI (i.e. the UK Science Museum, National Media Museum and National Railway Museum) back in March 2011.

There's more in the post I wrote on the museum developers blog at the time, Collections data published, but in summary:
We’ve released the files [218,822 object records, 40,596 media records and 173 event records] as a lightweight experiment – we’d like to understand whether, and if so, how, people would use our data. We’d also like to explore the benefits for the museum and for programmers using our data – your feedback will inform decisions about future investment in more structured data as well as helping shape our understanding of the requirements of those users. The files are in CSV format – because it’s a really simple format, viewable in a text editor, we hope that it will be usable by most people.
And since someone asked for some background on how I dealt with the organisational issues, the short answer is - I was pragmatic, figured any reasonable data was better than none, and kept it simple.  Or, as I wrote at the time in Update on collections data and geocoded NRM data:
A few people have commented on the licence (Creative Commons Attribution-NonCommercial-ShareAlike, CC BY-NC-SA) and on the format (CSV).  As tomorrow is my last day, I can’t really speak for the museum but the intention is to learn from how people use the data – the things they make, the barriers they face, etc – and iterate (as resources allow) until we get to an optimal solution (or solutions). So please get in touch if you’ve got requests or think you can help clear up some of the issues these kinds of projects face, because there’s a good chance you’ll help make a difference.

The licence is a pragmatic solution – it’s clarification of existing terms rather than a change to our terms, because this avoided a need for legal advice, policy review, etc, that would have added several months to the process.

And yes, I know CSV is quick and dirty, but it’s effective. The museum sector is still working out how to match the resources available with the needs of mash-up type developers who work best with JSON and those who are aiming for linked open data; my hope is that your feedback on this will help museums figure out how to support people using open data in various forms. A simple solution like this also means it’s easy for the museum to re-run the export to update the data as time goes on, and that anyone, geek or not, can open the files without being startled by angle brackets and acronyms. Also, did I mention it was quick?
In some ways, 2011 has been the year I really understood how much of a barrier a 'non-commercial' license is to re-use ('Wired releases images via Creative Commons, but reopens a debate on what “noncommercial” means' is quite a useful article for understanding the confusion though the LOD-LAM Summit was really where it came together for me).  Even I've struggled with questions like 'does a non-commercial license mean I can or can't upload the data to Google Fusion Tables to clean it?', let alone 'can a widget made with non-commercial data be displayed on an ad-supported blog site?'.

Most people who want to play with heritage data want to do the right thing, so an ambiguous 'non-commercial' license effectively prevents them using it (people who want to do bad things with it would probably just scrape the data anyway).  I get the sense that museums (and other GLAM orgs) are strongly loss averse, so a full 'commercial use ok' statement might be a bit much, but maybe we can do more to define exactly what's reasonable 'commercial' use and what's not?  The Wired article provides some useful starting questions, as does Europeana's discussion of their Data Exchange Agreement. Maybe 2012 will be the year we start to provide answers...

Update, January 2013: I've been writing a piece on open cultural data in museums so have been coming across more material on confusion about 'non-commercial'.  The Danger of Using Creative Commons Flickr Photos in Presentations discusses one case where the owner of a photograph was confused about whether it was being used commercially or not.  While that may turn out to be a case of mistaken identity, one commenter, Michael, says:
'Commercial and non-commercial are very difficult to determine. As such, I make a point of never using photos that have a non-commercial license. Too much hassle. (I also now do not use photos with a share-alike provision. Same reason, too much hassle.)'

A post on the Creative Commons blog, Library catalog metadata: Open licensing or public domain? discusses the case for and against requesting vs requiring attribution.

Saturday, 11 June 2011

'Share What You See' at hack4europe London

A quick report from hack4europe London, one of four hackathons organised by Europeana to 'showcase the potential of the API usage for data providers, partners and end-users'.

I have to confess that when I arrived I wasn't feeling terribly inspired - it's been a long month and I wasn't sure what I could get done at a one-day hack.  I was intrigued by the idea of 'stealth culture' - putting cultural content out there for people to find, whether or not they were intentionally looking for 'a cultural experience' - but I couldn't think of a hack about it I could finish in about six hours.  But I happened to walk past Owen Stephen's (@ostephens) screen and noticed that he was googling something about WordPress, and since I've done quite a lot of work in WordPress, I asked what his plans were.  After a chat we decided to work together on a WordPress plugin to help people blog about cool things they found on museum visits.  I'd met Owen at OpenCulture 2011 the day before (though we'd already been following each other on twitter) but without the hackday it's unlikely we would have ever worked together.

So what did we make?  'Share What You See' is a plugin designed to make a museum and gallery visit more personal, memorable and sociable.  There's always that one object that made you laugh, reminded you of friends or family, or was just really striking.  The plugin lets you search for the object in the Europeana collection (by title, and hopefully by venue or accession number), and instantly create a blog post about it (screenshot below) to share it with others.
Screenshot: post pre-populated with information about the object. 
Once you've found your object, the plugin automatically inserts an image of it, plus the title, description and venue name.

You can then add your own text and whatever other media you like.  The  plugin stores the originally retrieved information in custom fields so it's always there for reference if it's updated in the post.  Once an image or other media item is added, you can use all the usual WordPress tools to edit it.

If you're in a gallery with wifi, you could create a post and share an object then and there, because WordPress is optimised for mobile devices.  This help makes collection objects into 'social objects', embedding them in the lives of museum and gallery visitors.  The plugin could also be used by teachers or community groups to elicit personal memories or creative stories before or after museum visits.

The code is at https://github.com/mialondon/Share-what-you-see and there's a sample blog post at http://www.museumgames.org.uk/jug/.  There's still lots of tweaks we could have made, particularly around dealing with some of the data inconsistencies, and I'd love a search by city (in case you can't quite remember the name of the museum), etc, but it's not bad for a couple of hours work and it was a lot of fun.  Thanks to the British Library for hosting the day (and the drinks afterwards), the Collections Trust/Culture Grid for organising, and Europeana for setting it up, and of course to Owen for working with me.  Oh, and we won the prize for "developer's choice" so thank you to all the other developers!

Sunday, 16 January 2011

Notes from Culture Hack Day (#chd11)

Culture Hack Day (#chd11) was organised by the Royal Opera House (the team being @rachelcoldicutt, @katybeale, @beyongolia, @mildlydiverting, @dracos - and congratulations to them all on an excellent event). As well as a hack event running over two days, they had a session of five minute 'lightning talks' on Saturday, with generous time for discussion between sessions. This worked quite well for providing an entry point to the event for the non-technical, and some interesting discussion resulted from it. My notes are particularly rough this time as I have one arm in a sling and typing my hand-written notes is slow.

Lightning Talks
Tom Uglow @tomux “What if the Web is a Fad?”
'We're good at managing data but not yet good at turning it into things that are more than points of data.' The future is about physical world, making things real and touchable.

Clare Reddington, @clarered, “What if We Forget about Screens and Make Real Things?”
Some ace examples of real things: Dream Director; Nuage Vert (Helsinki power station projected power consumption of city onto smoke from station - changed people's behaviour through ambient augmentation of the city); Tweeture (a conch, 'permission object' designed to get people looking up from their screens, start conversations); National Vending Machine from Dutch museum.

Leila Johnston, @finalbullet talked about why the world is already fun, and looking at the world with fresh eyes. Chromaroma made Oyster cards into toys, playing with our digital footprint.

Discussion kicked off by Simon Jenkins about helping people get it (benefits of open data etc) - CR - it's about organisational change, fears about transparency, directors don't come to events like this. Understand what's meant by value - cultural and social as well as economic. Don't forget audiences, it has to be meaningful for the people we're making it (cultural products) for'.

Comment from @fidotheCultural heritage orgs have been screwed over by software companies. There's a disconnect between beautiful hacks around the edges and things that make people's lives easier. [Yes! People who work in cultural heritage orgs often have to deal with clunky tools, difficult or vendor-dependent data export proccesses, agencies that over-promise and under-deliver. In my experience, cultural orgs don't usually have internal skills for scoping and procuring software or selecting agencies so of course they get screwed over.]

TU: desire to be tangible is becoming more prevalent, data to enhance human experience, the relationship between culture and the way we live our lives.

CR: don't spend the rest of the afternoon reinforcing silos, shouldn't be a dichotomy between cultural heritage people and technologists. [Quick plug for http://museum30.ning.com/, http://groups.google.com/group/antiquist, http://museum-api.pbwiki.com/ and http://museumscomputergroup.org.uk/email-list/ as places where people interested in intersection between cultural heritage and technology can mingle - please let me know of any others!] Mutual respect is required.

Tom Armitage, @infovore “Sod big data and mashups: why not hack on making art?”
Making culture is more important than using it. 3 trends: 1) collection - tools to slice and dice across time or themes; 2) magic materials 3) mechanical art, displays the shape of the original content; 3a) satire - @kanyejordan 'a joke so good a machine could make it'.

Tom Dunbar, @willyouhelp - story-telling possibilites of metadata embedded in media e.g. video [check out Waisda? for game designed to get metdata added to audio-visual archives]. Metadata could be actors, characters, props, action...

Discussion [?]:remixing in itself isn't always interesting. Skillful appropriation across formats... Universe of editors, filterers, not only creators. 'in editing you end up making new things'.

Matthew Somerville, @dracos, Theatricalia, “What if You Never Needed to Miss a Show?”
'Quite selfish', makes things he needs. Wants not to miss theatre productions with people he likes in/working on them. Theatricalia also collects stories about productions. [But in discussion it came up that the National Theatre asked him to remove data - why?! A recommendation system would definitely get me seeing more theatre, and I say that as a fairly regular but uninformed theatre-goer who relies on word-of-mouth to decide where to spend ticket money.]

Nick Harkaway, @Harkaway on IP and privacy
IP as way of ringfencing intangible ideas, requiing consent to use. Privacy is the same. Not exciting, kind of annoying but need to find ways to make it work more smoothly while still proving protection. 'Buying is voting', if you buy from Tesco, you are endorsing their policies. 'Code for the change you want to see in the world', build the tools you want cultural orgs to have so they can do better. [Update: Nick has posted his own notes at Notes from Culture Hack Day. I really liked the way he brought ethical considerations to hack enthusiasm for pushing the boundaries of what's possible - the ability to say 'no' is important even if a pain for others.]

Chris Thorpe, @jaggeree. ArtFinder, “What if you could see through the walls of every museum and something could tell you if you’d like it?”

Culture for people who don't know much about culture. Cultural buildings obscure the content inside, stop people being surprised by what's available. It's hard if you don't know where to start. Go for user-centric information. Government Art Collection Explorer - ace! Wants an angel for art galleries to whisper information about the art in his ear. Wants people to look at the art, not the screen of their device [museums also have this concern]. SAP - situated audio platform. Wants a 'flight data recorder' for trips around cultural places.

Discussion around causes of fear and resistance to open data - what do cultural orgs fear and how can they learn more and relax? Fear of loss of provenance - response was that for developers displaying provenance alongside the data gives it credibility; counter-response was that organisations don't realise that's possible. [My view is that the easiest way to get this to change is to change the metrics by which cultural heritage organisations are judged, and resolve the tension between demands to commercialise content to supplement government grants and demands for open access to that same data. Many museums have developed hybrid 'free tombstone, low-res, paid-for high-res' models to deal with this, but it's taken years of negotiation in each institution.] I also ranted about some of these issues at OpenTech 2010, notes at 'Museums meet the 21st century'.

Other discussion and notes from twitter - re soap/drama characters tweeting - I managed to out myself as a Neighbours watcher but it was worth it to share that Neighbours characters tweet and use Facebook. Facebook relationship status updates and events have been included as plot points, and references are made to twitter but not to the accounts of the characters active on the service. I wonder if it's script writers or marketing people who write the characters tweets? They also tweet in sync with the Australian showings, which raises issues around spoilers and international viewers.

Someone said 'people don't want to interact with cultural institutions online. They want to interact with their content' but I think that's really dependent on the definition of content - as pointed out, points of data have limited utility without further context. There's a catch-22 between cultural orgs not yet making really engaging data and audiences not yet demanding it, hopefully hack days like CHD11 help bridge the gap and turn data into stories and other meaningful content. We're coming up against the limits of what can be dome programmatically, especially given variation in quality and extent of cultural heritage data (and most of it is data rather than content).

[Update: after writing this I found a post The lightning talks at Culture Hack Day about the day, which happily picks up on lots of bits I missed. Oh, and another, by Roo Reynolds.]

After the lightning talks I popped over the road to check out the hacking and ended up getting sucked in (the lure of free pizza had a powerful effect!).  I worked on a WordPress plugin with Ian Ibbotson @ianibbo that lets you search for a term on the Culture Grid repository and imports the resulting objects into my museum metadata games so that you can play with objects based on your favourite topic.  I've put the code on github [https://github.com/mialondon/mmg-import] and will move it from my staging server to live over the next few days so people can play with the objects.  It's such a pain only having one hand, and I'm very grateful to Ian for the chance to work together and actually get some code written.  This work means that any organisation that's contributed records to the Culture Grid can start to get back tags or facts to enhance their collections, based on data generated by people playing the games.  The current 300-ish objects have about 4400 tags and 30 facts, so that's not bad for a freebie. OTOH, I don't know of many museums with the ability to display content created by others on their collections pages or store it in their collections management systems - something for another hack day?

Something I think I'll play around with a bit more is the idea of giving cultural heritage data a quality rating as it's ingested.  We discussed whether the ratings would be local to an app (as they could be based on the particular requirements of that application) or generalised and recorded in the CultureGrid service.  You could record the provence of a rating which might be an approach that combines the benefits of both approaches.  At the moment, my requirements for a 'high quality' record would be: title (e.g. 'The Ashes trophy', if the object has one), name or type of object (e.g. cup), date, place, decent sized image, description.

Finally, if you're interested in hacking around cultural heritage data, there's also historyhackday next weekend. I'm hoping to pop in (dependent on fracture and MSc dissertation), not least because in March I'm starting a PhD in digital humanities, looking at participatory digitisation of geo-located historical material (i.e. getting people to share the transcriptions and other snippets of ad hoc digitisation they do as part of their research) and it's all hugely relevant.

Wednesday, 3 November 2010

What would Phar Lap do? AKA, what happens when Facebook and museum URIs meet a dead horse?

Phar Lap was a famous race horse.  After he died (in film-worthy suspicious circumstances), bits of Phar Lap ended up in three different museums - his skin is at Melbourne Museum, his skeleton is at Te Papa in Wellington, NZ, and his heart is in Canberra at the National Museum of Australia.

I've always been fascinated by the way the public respond to Phar Lap - when I worked at Museum Victoria, the outreach team would regularly get emails written to Phar Lap by people who had seen the film or somehow come across his story.  (I was also never quite sure why they thought emailing a dead horse would work).  So when I first heard that Phar Lap was on Facebook, I was curious to see which museum would have 'claimed' Phar Lap.  Does possession of the most charismatic object (the hide) make it easier for Melbourne Museum to step up as the presence of Phar Lap on social media, or were they just the first to be in that space?  The issues around 'ownership' and right to speak for an iconic object like Phar Lap make a brilliant case study for how museums represent their collections online.

And today, when I came across three posts (Responses to "Progress on Museum URIs", Progress on Museum URIs by @sebastianheath, Identifing Objects in Museum Collections by @ekansa) on movements towards stable museum URIs that problematised the "politics of naming and identifying cultural heritage" and the concept of the "exclusive right of museums to identify their objects", I thought of Phar Lap.  (Which is nice, cos 80 years and one day ago he won the Melbourne Cup).

Of the three museums that own bits of the dead horse, which gets to publish the canonical digital record about Phar Lap? I hope the question sounds silly enough to highlight the challenges and opportunities in translating physical models to the digital realm. Of course each museum can publish a record (specifically, mint a URI) about Phar Lap (and I hope they do) but none of the museums could prevent the others from publishing (and hopefully they wouldn't want to).

Or as the various blog posts said, "many agents can assert an identity for an object, with those identities together forming a distributed and diverse commentary on the human past", and museums need to play their part: "a common identifier promoted by and discoverable at the holding institution will ease the process of recognizing that two or more identifiers refer to the 'same thing'".

Of course it's not that simple, and if you're interested in the questions the museum sector (by which I hopefully don't only mean me) is grappling with, the museums and the machine-processable web page on Permanent IDs has links to discussions on the MCG list, and I've wrestled a bit with how URIs might look at the Science Museum/NMSI (and I need to go back and review the comments left by various generous people).  I'd love to know what other museums are planning, and what consumers of the data might need, so that we can come up with a robust common model for museum URIs.

And to reward you for getting this far, here is a picture of Phar Lap on Facebook as his skin and bones are about to be re-united:

Sunday, 24 October 2010

UK Culture Grid wants to know what developers need - get in!

Neil Smith from Knowledge Integration dropped by the Museums and the machine-processable web wiki to ask what users (developers) need to get data in and out of the Culture Grid:
To support the ambitious targets for increasing the number of item records in Culture Grid, we thought know would be a good time to review the venerable old application profile we use for importing metadata into the Grid. I've added a discussion page reviewing options at http://museum-api.pbworks.com/w/page/Culture-Grid-Profile.

We really want the community to be involved in helping ensure that whatever profile (or profiles) we support will meet the needs of users - not only for getting things into the grid but also for getting things out in a format that is useful to them. Although the paper focusses mainly on XML representations of metadata, we're also interested in your views on whether non-XML representations (e.g RDF or JSON) need to be supported.
So whether you work in a museum or are an external developer who'd like to use museum data, I'd encourage you to think about the four options Neil outlines, and to comment, ask questions, share sample data, vote for your favourite option, whatever, on the Culture Grid Profile page.  One of the options is to develop a new model - definitely more time-consuming, but a great opportunity to make your needs known.

As an indication of the type of content that's available through the Culture Grid, I've copied this text from some of their about pages: "It contains over 1 million records from over 50 UK collections, covering a huge range of topics and periods.  Records mostly refer to images but also text, audio and video resources and are mostly about museum objects with library, archive and other kinds of collections also included."  So, that's:

  • "information about items in collections (referencing the images, video, audio or other material you offer online about the things in your collections)
  • information about collections as a whole (their scope, significance and access details)
  • information about collecting organisations (contact and access details)"

There's a lot of cultural heritage and tech jargon involved on the Culture Grid Profile discussion page - don't hold back on asking for clarifications where needed.  I'm certainly not an expert on the various schemas and it's a very long time since I helped work out the Exploring 20th Century London extensions for the original PNDS, but I've given it a go.

If you've read this far, you might also be interested in the first ever Culture Grid Hack Day in Newcastle Upon Tyne on December 3, 2010.

Sunday, 12 September 2010

'Museums meet the 21st century' - OpenTech 2010 talk

These are my notes for the talk I gave at OpenTech 2010 on the subject of 'Museums meet the 21st Century'. Some of it was based on the paper I wrote for Museums and the Web 2010 about the 'Cosmic Collections' mashup competition, but it also gave me a chance to reflect on bigger questions: so we've got some APIs and we're working on structured, open data - now what? Writing the talk helped me crystallise two thoughts that had been floating around my mind. One, that while "the coolest thing to do with your data will be thought of by someone else", that doesn't mean they'll know how to build it - developers are a vital link between museum APIs, linked data, etc and the general public; two, that we really need either aggregated datasets or data using shared standards to get the network effect that will enable the benefits of machine-readable museum data. The network effect would also make it easier to bridge gaps in collections, reuniting objects held in different institutions. I've copied my text below, slides are embedded at the bottom if you'd rather just look at the pictures. I had some brilliant questions from the audience and afterwards, I hope I was able to do them justice. OpenTech itself was a brilliant day full of friendly, inspiring people - if you can possibly go next year then do!

Museums meet the 21st century.
Open Tech, London, September 11, 2010

Hi, I'm Mia, I work for the Science Museum, but I'm mostly here in a personal capacity...

Alternative titles for this talk included: '18th century institution WLTM 21st century for mutual benefit, good times'; 'the Age of Enlightenment meets the Age of Participation'. The common theme behind them is that museums are old, slow-moving institutions with their roots in a different era.

Why am I here?

The proposal I submitted for this was 'Museums collaborating with the public - new opportunities for engagement?', which was something of a straw man, because I really want the answer to be 'yes, new opportunities for engagement'. But I didn't just mean any 'public', I meant specifically a public made up of people like you. I want to help museums open up data so more people can access it in more forms, but most people can't just have a bit of a tinker and create a mashup. “The coolest thing to do with your data will be thought of by someone else” - but that doesn’t mean they’ll know how to build it. Audiences out there need people like you to make websites and mobile apps and other ways for them to access museum content - developers are a vital link in the connection between museum data and the general public.

So there's that kind of help - helping the general public get into our data; and there's another kind of help - helping museums get their data out. For the first, I think I mostly just want you to know that there's data out there, and that we'd love you to do stuff with it.

The second is a request for help working on things that matter. Linkable, open data seems like a no-brainer, but museums need some help getting there.

Museums struggle with the why, with the how, and increasingly with the "we are reducing our opening hours, you have to be kidding me".

Chicken and the egg

Which comes first - museums get together and release interesting data in a usable form under a useful licence and developers use it to make cool things, or developers knock on the doors of museums saying 'we want to make cool things with your data' and museums get it sorted?

At the moment it's a bit of both, but the efforts of people in museums aren't always aligned with the requests from developers, and developers' requests don't always get sent to someone who'll know what to do with it.

So I'm here to talk about some stuff that's going on already and ask for a reality check - is this an idea worth pursuing? And if it is, then what next?
If there’s no demand for it, it won’t happen. Nick Poole, Chief Executive, Collections Trust, said on the Museums Computer Group email discussion list: "most museum people I speak to tend not to prioritise aggregation and open interoperability because there is not yet a clear use case for it, nor are there enough aggregators with enough critical mass to justify it.”

But first, an example...

An experiment - Cosmic Collections, the first museum mashup competition

The Cosmic Collections project was based on a simple idea - what if a museum gave people the ability to make their own collection website for the general public? Way back in December 2008 I discovered that the Science Museum was planning an exhibition on astronomy and culture, to be called ‘Cosmos & Culture’. They had limited time and resources to produce a site to support the exhibition and risked creating ‘just another exhibition microsite’. I went to the curator, Alison Boyle, with a proposal - what if we provided access to the machine-readable exhibition content that was already being gathered internally, and threw it open to the public to make websites with it? And what if we motivated them to enter by offering competition prizes? Competition participants could win a prize and kudos, and museum audiences might get a much more interesting, innovative site. Astronomy is one of the few areas where the amateur can still make valued scientific contributions, so the idea was a good match for museum mission, exhibition content, technical context, and hopefully developers - but was that enough?

The project gave me a chance to investigate some specific questions. At the time, there were lots of calls from some quarters for museums to produce APIs for each project, but there was also doubt about whether anyone would actually use a museum API, whether we could justify an investment in APIs and machine-readable data. And can you really crowdsource the creation of collections interfaces? The Cosmic Collections competition was a way of finding out.

Lessons? An API isn't a magic bullet, you still need to support the dev community, and encourage non-technical people to find ways to play with it. But the project was definitely worth doing, even if just for the fact that it was done and the world didn't end. Plus, the results were good, and it reinforced the value of working with geeks. [It also got positive coverage in the technical press. Who wouldn’t be happy to hear ‘the museum itself has become an example of technological innovation’ or that it was ‘bringing museums out into the open as places of innovation’?]

Back to the chicken and the egg - linking museums

So, back to the chicken and the egg... Progress is being made, but it gets bogged down in discussions about how exactly to get data online. Museums have enough trouble getting the suppliers they work with to produce code that meets accessibility standards, let alone beautifully structured, re-usable open data.

One of the reasons open, structured data is so attractive to museum technologists is that we know we can never build interfaces to meet the needs of every type of audience. Machine-readable data should allow people with particular needs to create something that supports their own requirements or combines their data with ours to make lovely new things.

Explore with us - tell museums what you need

So if you're someone who wants to build something, I want to hear from you about what standards you're already working with, which formats work best for you...

To an extent that's just moving the problem further down the line, because I've discovered that when you ask people what data standards they want to use, and they tell you it turns out they're all different... but at least progress is being made.

Dragons we have faced

I think museums are getting to the point where they can live with the 80% in the interest of actually getting stuff done.

Museums need to get over the idea that linkable data must be perfect - perfectly clean data, perfectly mapped to perfect vocabularies and perfectly delivered through perfect standards. Museums are used to mapping data from their collections management systems for a known end-use, they've struggled with open-ended requirements for unknown future uses.

The idea that aggregated data must be able to do everything that data provided at source can do has held us back. Aggregated data doesn't need to be able to do everything - sometimes discoverability is enough, as long as you can get back to the source if you need the rest of the data. Sometimes it's enough to be able to link to someone else's record that you've discovered.

Museum data and the network effect

One reason I'm here (despite the fact that public speaking is terrifying) is a vision of the network effect that could apply when we have open museum data.

We could re-unite objects across time and place and people, connecting visitors and objects, regardless of owing institution or what type of object or information it is. We could create highlight collections by mining data across museums, using the links people are making between our collections. We can help people tell their local stories as well as the stories about big subject and world histories. Shared data standards should reduce learning curve for people using our data which would hopefully increase re-use.

Mismatches between museums and tech - reasons to be patient

So that's all very exciting, but since I've also learnt that talking about something creates expectations, here are some reasons to be patient with museums, and tolerant when we fail to get it right the first time...

IT is not a priority for most museums, keeping our objects secure and in one piece is, as is getting some of them on display in ways that make sense to our audiences.

Museums are slow. We'll be talking about stuff for a long time before it happens, because we have limited resources and risk-averse institutions. Museum project management is designed for large infrastructure projects, moving hundreds of delicate objects around while major architectural builds go on. It's difficult to find space for agility and experimentation within that.

Nancy Proctor from the Smithsonian said this week: "[Museum] work is more constrained than a general developer" - it must be of the highest quality; for everybody - public good requires relevance and service for all, and because museums are in the 'forever business' it must be sustainable.

How you can make a difference

Museums are slowly adapting to the participation models of social media. You can help museums create (backend) architectures of participation. Here are some places where you can join in conversations with museum technologists:

Museums Computer Group – events, mailing list http://museumscomputergroup.org.uk/ #ukmcg @ukmcg

Linking Museums – meetups, practical examples, experimenting with machine-readable data http://museum-api.pbworks.com/

Space Time Camp - Nov 4/5, #spacetimecamp

‘Museums and the Web’ conference papers online provide a good overview of current work in the sector http://www.archimuse.com/conferences/mw.html

So that‘s all fun, but to conclude - this is all about getting museums to the point where the technology just works, data flows like water and our energy is focussed on the compelling stories museums can tell with the public. If you want to work on things that matter - museums matter, and they belong to all of us - we should all be able to tell stories with and through museums.

Thank you for listening


Keep in touch at @mia_out or http://openobjects.blogspot.com/

Thursday, 29 April 2010

Slides and talk from 'Cosmic Collections' paper

This is a lazy post, a straight copy and paste of my presentation notes (my excuse is that I'm eight days behind on everything at work and uni after being grounded in the US by volcanic ash). Anyway, I hope you enjoy it or that it's useful in some way.
Cosmic Collections: creating a big bang?
View more presentations from Mia .

Slide 1 (solar rays - Cosmic Collections):
The Cosmic Collections project was based on a simple idea - what if we gave people the ability to make their own collection website? The Science Museum was planning an exhibition on astronomy and culture, to be called ‘Cosmos & Culture’. We had limited time and resources to produce a site to support the exhibition and we risked creating ‘just another exhibition microsite’. So what if we provided access to the machine-readable exhibition content that was already being gathered internally, and threw it open to the public to make websites with it?  And what if we motivated them to enter by offering competition prizes?  Competition participants could win a prize and kudos, and museum audiences might get a much more interesting, innovative site.
The idea was a good match for museum mission, exhibition content, technical context, hopefully audience - but was that enough?

Slide 2 (satellite dish):
Questions...
If we built an API, would anyone use it?
Can you really crowdsource the creation of collections interfaces?

The project gave me a chance to investigate some specific questions.  At the time, there were lots of calls from some quarters for museums to produce APIs for each project, but would anyone actually use a museum API?  The competition might help us understand whether or how we should invest in APIs and machine-readable data.
We can never build interfaces to meet the needs of every type of audience.  One of the promises of machine-readable data is that anyone can make something with your data, allowing people with particular needs to create something that supports their own requirements or combines their data with ours - but would anyone actually do it?

Slide 3 (map mashup):
Mashups combine data from one or more sources and/or data and visualisation tools such as maps or timelines.

I'm going to get the geek stuff out of the way and quickly define mashups and APIs...
Mashups are computer applications that take existing information from known sources and present it to the viewer in a new way. Here’s a mashup of content edits from Wikipedia with a map showing the location of the edit.

Slide 4 (APIs)
APIs (Application Programming Interfaces) are a way for one machine to talk to another: ‘Hi Bob, I’d like a list of objects from you, and hey, Alice, could you draw me a timeline to put the objects on?’

APIs tell a computer, 'if you go here, you will get that information, presented like this, and you can do that with it'.
A way of providing re-usable content to the public, other museums and other departments within our museum - we created a shared backend for web and gallery interactives.
I think of APIs as user interfaces for developers and wanted to design a good experience for developers with the same care you would for end users*.  I hoped that feedback from the competition could be used to improve the beta API

* we didn’t succeed in the first go but it’s something to aim for post-beta

Slide 5: (what if nobody came?)
AKA 'the fears and how to deal with them'
Acknowledge those fears
Plan for the worst case scenario
Take a deep breath and do it anyway

And on the next slides, the results.  If I was replicating the real experience, you’d have several nerve-biting months while you waited for the museum to lumber into gear, planned the launch event, publicised the project in the participant communities... Then waited for results to come in. But let’s skip that bit...

Slide 6: (Ryan Ludwig's http://www.serostar.com/cosmic/)
The results - our judges declared a winner and a runner-up, these are screenshots - this is the second prize winning entry.
People came to the party. Yay! I'd like to thank all the participants, whether they submitted a final entry or not. It wouldn't have worked without them.

Slide 7: (Natalie and Simon's http://cosmos.natimon.com/)
This is a screenshot from the winning site - it made the best use of the API and was designed to lure the visitor in and keep drawing them through the site.

(We didn’t get subject specialists scratching their own itch - maybe they don’t need to share their work, maybe we didn’t reach them. Would like to reach researchers, let them know we have resources to be used, also that they can help us/our audiences by sharing their work)

Slide 8: (astrolabe - what did we learn?)
People need (more) help to participate in a geektastic project like this
The dynamics of a competition are tricky
Mashups are shaped by the data provided - you get out what you put in

Can we help people bring their own content to a future mashup?

Slide 9: (evaluation)
I did a small survey to evaluate the project... Turns out the project was excellent outreach into the developer community. People were really excited about being invited to play with our data.  My favourite quote: "The very idea of the competition was awesome"

Slide 10: (paper sheet)
Also positive coverage in technical press. So in conclusion?

Slide 11: (Tim Berners-Lee):
“The thing people are amazed about with the web is that, when you put something online, you don’t know who is going to use it—but it does get used.”

There are a lot of opportunities and excitement around putting machine-readable data online...

Slide 12: Tim Berners-Lee 2:
But:  It doesn’t happen automatically; It’s not a magic bullet

But people won't find and use your APIs without some encouragement. You need to support your API users. People outside the museum bring new ideas but there's still a big role for people who really understand the data and audiences to help make it a quality experience...

Slide 13 (space):
What next?
Using the feedback to focus and improve collection-wide API
Adding other forms of machine-readable data
Connecting with data from your collections?

I've been thinking about how to improve APIs - offer subject authorities with links to collections, embed markup in the collections pages to help search engines understand our data...
I want more! The more of us with machine-readable data available for re-use, the better the cross-collections searches, the region or specialism-wide mashups... I'd love to be able to put together a mashup showing all the cultural heritage content about my suburb; all the Boucher self-portraits; all the inventions that helped make the Space Shuttle work...

Slide 14: (thank you)
If you're interested in possibilities of machine-readable data and access to your collections, join in the conversation on the museum API wiki or follow along on twitter or on blogs.  Join in at http://museum-api.pbworks.com/
More at http://openobjects.blogspot.com/ or @mia_out

Full paper online at http://www.archimuse.com/mw2010/papers/ridge/ridge.html

Image credits include:
http://antwrp.gsfc.nasa.gov/apod/ap100415.html
http://antwrp.gsfc.nasa.gov/apod/ap100414.html
http://antwrp.gsfc.nasa.gov/apod/ap100409.html
http://antwrp.gsfc.nasa.gov/apod/ap100209.html
http://antwrp.gsfc.nasa.gov/apod/ap100315.html
http://www.sciencemuseum.org.uk/Centenary/Home/Icons/Pilot_ACE_Computer.aspx
http://www.prospectmagazine.co.uk/2010/01/mash-the-state/

Sunday, 4 April 2010

'Cosmic Collections' - my MW2010 paper online

My Museums and the Web 2010 paper is up at Cosmic Collections: Creating a Big Bang and I'm working on the slides now and I'm curious - what would you like to see more of in a presentation?  It's only short (6 minutes) so I'm currently thinking setup (including lots of definitions for non-geeks), outcomes (did the project succeed?), and a bit on what I think the next steps are (basically a call to get your data online in re-usable formats).

I'm thinking of leading with this Tim Berners-Lee quote from an article in Prospect, Mash the state:
"The thing people are amazed about with the web is that, when you put something online, you don't know who is going to use it—but it does get used."

Thursday, 21 January 2010

Cosmic Collections - the results are in. And can you help us ask the right questions?

For various reasons, the announcement of the winners of our mashup competition has been a bit low key - but we're working on a site that combines the best bits of the winners, and we'll make a bit more of a song and dance about it when that's ready.

I'd like to take the opportunity to personally thank the winners - Simon Willison and Natalie Down in first place, and Ryan Ludwig as runner-up - and equally importantly, those who took part but didn't win; those who had a play and gave us some feedback; those who helped spread the word, and those who cheered along the way.

I have a cheeky final request for your time.  I would normally do a few interviews to get an idea of useful questions for a survey, but it's not been possible lately. I particularly want to get a sense of the right questions to ask in an evaluation because it's been such a tricky project to explain and 'market', and I'm far too close to it to have any perspective.  So if you'd like to help us understand what questions to ask in evaluation, please take our short survey http://www.surveymonkey.com/s/5ZNSCQ6 - or leave a comment here or on the Cosmic Collections wiki.  I'm writing a paper on it at the moment, so hopefully other museums (and also the Science Museum itself) will get to learn from our experiences.

And again - my thanks to those who've already taken the survey - it's been immensely useful, and I really appreciate your honesty and time.

Thursday, 19 November 2009

Nine days to go! And entering Cosmic Collections just got easier

Quoting myself over on the museum developers blog, Cosmic Collections – do one thing and do it well:
I’ve realised that there may be some mismatch between the way mashups tend to work, and the scope we’ve suggested for entries to our competition. The types of interfaces someone might produce with the API may lend themselves more to exploring one particular idea in depth than produce something suitable for the broadest range of our audiences.
So I’m proposing to change the scope for entries to the competition, to make it more realistic and a better experience for entrants: I’d like to ask you to build a section of a site, rather than a whole site. The scope for entrants would then be: “create something that does one thing, and does it well”. Our criteria – use of collections data, creativity, accessibility, user experience and ease of deployment and maintenance – are still important but we’ll consider them alongside the type of mashup you submit.
I've updated the Cosmic Collections competition page to reflect this change. This page also features a new 'how to take part' section, including a direct link to the API and to a discussion group.

I'd love to hear your thoughts on this change - there's an email address lurking on the competition page, and I'm on twitter @mia_out and @coscultcom.

In other news, programmableweb published a blog post about the competition today: Science Museum Opens API and Challenges Developers to Mashup the Cosmos. Woo!

And I don't know if it's any kind of consolation if you're entering, but I'll be working right alongside you up until Friday 28th, on an assignment for my MSc.

Thursday, 22 October 2009

'Cosmic Collections' launches at the Science Museum this weekend

I think I've already said pretty much everything I can about the museum website mashup competition we're launching around the 'Cosmos and Culture' exhibition, but it'd be a bit silly of me not to mention it here since the existence and design of the project reflects a lot of the issues I've written about here.

If you make it along to the launch at the Science Museum on Saturday, make sure you say hello - I should be easy to find cos I'm giving a quick talk at some point.

Right now the laziest thing I could do is to give you a list of places where you can find out more:

Finally, you can talk to us @coscultcom on twitter, or tag content with #coscultcom.

Btw - if you want an idea of how slowly museums move, I think I first came up with the idea in January (certainly before dev8D because it was one of the reasons I wanted to go) and first blogged about it (I think) on the museum developers blog in March. The timing was affected by other issues, but still - it's a different pace of life!

Sunday, 21 June 2009

Museum pecha kucha night

The first museum pecha kucha night was held in London at the British Museum on June 18, 2009. I took rough notes during the presentations, and have included the slides and notes from my own presentation. The event used the tag 'mwpkn' to gather together tweets, photos, etc. The focus of this first museum pecha kucha was on sharing insights and inspiration from the Museums and the Web conference held in Indianapolis in April.

The event was organised by Shelley Mannion, who introduced the event, emphasising that it was about fun and connecting the museum tech community in an interesting way.

Gail Durbin (V&A), takeaways from MW2009
She's a practical person, looks for ideas to nick. Good idea as things get hazy after a conference, good intentions disappear.

First takeaway - Dina Helal let her play with her iPhone, decided she had to have one. She liked her mobile for the first time in her life.

Second - twittering was very important. Decided to do something with it. Twittering is hard, sending out messages that are interesting is difficult.

Enthusiasm at conferences is short lived - e.g. people excited about wedding site, but did they send in wedding photos? She talked to people about a self-portraiture idea, 'life on a postcard', but hasn't had a single response.

RSS feeds - came away knowing we had to review our RSS feeds, had been without attention for a long time.

Learnt that wikis are very hard work, they don't automatically look after themselves.
Creative use of Flickr - museum 'my karsh' collection

Resolved that had to work with Development. Looking at something like the British Library's - adopt a book for fathers day.

Something that bothers her - many museums think of 'Web 2.0' just as more channels to push out information, there's no sense of pulling in information about visitors.

Beck Tench, one of the most interesting people she met at the conference - practice and work go together very closely. Flickr plant project. She wants to get staff involved - has meeting on Fridays, in local bar, tweets to everyone, conducts something called Experimonth.

Last thing learnt - librarians have better cakes.

Silvia Filippini Fantoni (British Museum and Sorbonne University)
Silvia makes a plea for extra seconds as a non-native speaker (and synthesis not the best feature of Italians). Lecturer in museum informatics and evaluation methods at Sorbonne and project manager for multimedia guide project at British Museum.

So her focus at the conference was mostly on guides. Particularly Samis and Pau and others. Mini workshops and workshops on the topic before and during the conference. Demos from Paul Clifford (Museum of London). Exhibitors. Lots of museums are planning to develop applications.

Interest in using mobile technology as an interpretive tool is constantly growing, especially delivered on visitors own devices. Proliferations of mobile platforms. Proliferation of different functionalities - not just audio - visual, games, way finding, web access and communication, notes and comments. Have all these new platforms and functionalities improved the visitor experience? Yes, but there are some disadvantages.

Asks: aren't we trying to do too much? Are we trying to turn a useful interpretive tool into something too complex? Aren't we forgetting about core audio guide audience?

Are people interested in using their own devices? Do they have the time to pre-download, do they bring their devices? Samis and Pau - the answer is no/not yet. For the medium and short term still need to provide media in the museums. Touch screen devices are easier to use. Limited functionality makes interface simpler. Focus on content - AV messages, touch and listen.
Importance of sharing and learning from best practice. Some efforts at and after MW2009 - handheldconference.org. Discussion of developing open source content management system for mobile devices - contact Nancy Proctor.

Daniel Incandela (Indianapolis Museum of Art)
He's from America so should have extra time too. Also sick and medicated (so at least one of us will have a good time during the presentation).

Enjoys robots, dinosaurs, football and a good point. On holiday while here.

Slide - Shelley's twitter profile - she's responsible for him being here while on holiday.

He blogged about preparing for the presentation and got a comment from one of the pecha kucha founders - the main thing is to have fun, be passionate about something you love.

Twitterfall on the big screen was a major breakthrough at MW2009, (#mw2009 trended as a topic and attracted the attention of) pantygirl.

Digital story telling and tech can't happen without support, Max Anderson has been dream leader.

He's here representing IMA so going to showcase some projects - Roman Art from Louvre webisodes - paved the way for informal, agile, multiple content source creation.

Art Babble. IMA blog - ripped off other museums - gives many departments from museum a digital voice.

Half time experiment with awkward silence (blank slide). [In the pub afterwards, I discovered that this actually made at least one of the English people feel socially awkward!]

Brooklyn Museum - for him the real innovators for digital content for museums, won many awards at MW2009.

Te Papa's 'build a squid' had him at 'hello'. First example of a museum project that actually went viral?

Perhaps we could upgrade MW site? Better integration of social media, multimedia from previous conferences.

Loves Bruce Wyman - reason to go to MW2010.

art:21 - smart team, good approaches to publishing across platforms.

Wonders about agility - love new and emerging projects (?) we hear about at conferences, but how do we face an idea and deal with own internal issues?

The Dutch at Indy (were great) - but somewhere outside north America next for Museums and the Web?

Philip Poole (British Museum)
Everything I got from MW2009 can be put into one statement - spread it about. Enable your content to be spread by other people through APIs.

Does spreading out content dilute our authority? By putting it onto other websites, putting it in contact with other people. No, of course not.

Video was big at MW2009.

If going to use different platforms, will people come? We need to tailor content to different websites - can't just build it and assume people will come. Persian coins vs. ritual Mayan sacrifice on YouTube - which will get bigger audience? [Pick content delivery to suit audience and context.]

Platforms include ArtBabble, YouTube (shorter, edgier), iTunes U. Viral content - we can put features on our website, but a YouTube or Vimeo audience are going to spread things better. iTunes, U, can download and listen on train - takes out of website entirely.

Stats are important - e.g. need to include stats of video on different platforms, make sure people above you recognise the value in that. DCMS - very basic stats - perhaps they should be asking for different stats. "If DCMS ask how much video we put on YouTube, we'd all start doing it." [Brilliant point]

API - take content from website and put elsewhere. IMA Explore section - advertise the repeating pattern in their URLs - someone used them but wasn't going very well, they got in contact with him and helped him succeed, now biggest referrer outside search engines. He wants to do that for the British Museum - he knows the quirks, the data.

Why the 'softly softly' approach? Creating an entire API interface is huge mountain, people above you will want to avoid it if you show them the size of the whole mountain.

Digital NZ - fantastic example. Can create custom search, embed on website, also into gallery and people can vote for it

The British Museum is a museum of the world for the world, why should their web presence be any different?

Mia Ridge (Science Museum)
Yes, that's me. My slides on 'Bubbles and Easter eggs - Museum Pecha Kucha' are on slideshare - scroll down the page for full text and notes - or available as a PDF (2mb).

I talked about:

  • keeping the post-conference momentum going, particularly the 'do one thing' idea;
  • museum technologists as 'double domain experts';
  • not hiding museum geeks like Easter eggs but making more of them as a resource;
  • the responsibilities of museum geeks as their expertise is recognised;
  • breaking down internal silos; intelligent failure;
  • broken metrics and better project design (pitch the goal, not the method);
  • audience expectations in 2009;
  • possible first questions for digital projects and taking a whole museum view for new projects;
  • who's talking/listening to your audiences? trust and respect your audiences;
  • your museum is an iceberg (lots of the good stuff is hidden);
  • (s)mash the system (hold a mashup day);
  • and a challenge for your museum - has the web fundamentally changed your organisation?

Frankie Roberto (Rattle)
Went to the conference with a 'fan' hat on, just really enjoys museums. Loved the zoo - live exhibits are interactive, visceral. Role of live interpretation - how could it work with digital technology? Everyone loves dinosaur - Indy Children's Museum. All museums should have a carousel (can't remember what he was going to say about it).

The Power of Children; making a difference - really powerful stories.

Still thinking about the idea of creating visceral experiences.

ArtBabble - shouldn't generally create silos but ArtBabble spotted that YouTube wasn't working for certain types of content.

Davis LAB - kiosks and sofa. Said 'we are on the web'.

Drupal - lots of museums switching to it.

Richard Morgan (V&A) on APIS - ask, what is your museum good at?, and build an API for that - it may not be collections stuff.

'Things to do' page on V&A. Good way of highlighting ways to interact on website.

Semantic data, Aaron's talk on interpretation of bias, relocation from Flickr photos.
Breaking down ideas about authority on where an area is bounded by. OpenStreetMap - wants to add a historical layer to that so can scroll backwards and forwards in time. [I should ask whether this means layering old maps (with older street layouts like pre-Great Fire of London, or earlier representations?). Geo-rectification is expensive because it's time-consuming, but could it be crowdsourced? Geo-locating old images would be easier for the average person to do.]


Open Plaques - alpha project.

Thinks we won't need to digitise in the future as stuff will be born digital (ha, as if! Though it depends where you draw the lines about the end of collections - in my imagination they're like that warehouse scene at the end of Indiana Jones and the Raiders of the Lost Arc and we won't run out of things to properly digitise any time soon. Still, it's a useful question.)

Dan Zambonini (Box UK)
'Every film needs a villain'. In his impressions and insights from MW2009 he'll say things we may or may not agree with.

Slide - stuff we can do vs. stuff we can't do on either side of a gulf of perceived complexity. It's hard to progress from one to the other. Three questions to bridge gap - how to make relevant to everyday job, how to show advantages, how to make it easy.

Then he realised should talk about personal things - people and connections made. About people, stuff that happens in the evening. The evening drinks don't happen at UKMW - it's a shame we have to go to the other side of the world to talk to each other. [It does it you're at an event like mashed museum the day before - another reason to open it up to educators, curators, etc.]

Small museums vs. big museums - [should make stuff accessible to small museums.] Can get value by helping people. (He tells his ex-girlfriend that ) small is the new big. Also small quick wins. Break down the big things into smaller things, find ways can get to them through small changes in behaviour, bits of information.

How small is small? Greater or less than one day. If less than a day, might as well try it. If it's going to take a week, not small.

Museums should share data - not just as API - share data on traffic, spill gossip on marketing costs, etc. [Information is power, etc]

Celebrate failure - admit that some things go wrong.

Bigger picture - be honest. Tell us when to shut up (on e.g. the MCG email list?). Sometimes feel like there's too much politics. I know some of us can be a bit hot-headed but it's frustration (not meanness).

"never willingly outsource creativity" - that's rubbish.

If not on twitter, get on it. The more people talking to each other, the more powerful we are as a group. [But what happens if you miss a few days of twitter? I like twitter, but it's inaccessible if you don't have time to constantly keep up, or don't have a computer at home. Still, getting more people talking is an excellentbl point, even if twitter itself doesn't work for some people.]

The sector is missing practical, specific blog, not news and opinions. [Do collections system specific user groups take the place of blogs?]

Use grants to innovate and produce open source stuff. Right now private agencies will take a lot of the strain of applying for grants.

Sort out that copyright stuff. How difficult can it be?

Final slide summing up and last bit of innuendo. 'Beer makes you more attractive' - it's the after sessions stuff at conferences that's so valuable.

Frankie, Dan and Daniel's slides are also available in the 'Museum Tech Pecha Kucha' event on slideshare (and mine has now got an audio track, thanks to Shelley).

Wednesday, 13 May 2009

Final thoughts on open hack day (and an imaginary curatr)

I think hack days are great - sure, 24 hours in one space is an artificial constraint, but the sheer brilliance of the ideas and the ingenuity of the implementations is inspiring. They're a reminder that good projects don't need to take years and involve twenty circles of sign-off, even if that's the reality you face when you get back to the office.

I went because it tied in really well with some work projects (like the museum metadata mashup competition we're running later in the year or the attempt to get a critical mass of vaguely compatible museum data available for re-use) and stuff I'm interested in personally (like modern bluestocking, my project for this summer - let me know if you want to help, or just add inspiring women to freebase).

I'm also interested in creating something like a Dopplr for museums - you tell it what you're interested in, and when you go on a trip it makes you a map and list of stuff you could see while you're in that city.

Like: I like Picasso, Islamic miniatures, city museums, free wine at contemporary art gallery openings, [etc]; am inspired by early feminist history; love hearing about lived moments in local history of the area I'll be staying in; I'm going to Barcelona.

The 'list of cultural heritage stuff I like' could be drawn from stuff you've bookmarked, exhibitions you've attended (or reviewed) or stuff favourited in a meta-museum site.

(I don't know what you'd call this - it's like a personal butlr or concierge who knows both your interests and your destinations - curatr?)

The talks on RDFa (and the earlier talk on YQL at the National Maritime Museum) have inspired me to pick a 'good enough' protocol, implement it, and see if I can bring in links to similar objects in other museum collections. I need to think about the best way to document any mapping I do between taxonomies, ontologies, vocabularies (all the museumy 'ies') and different API functions or schemas, but I figure the museum API wiki is a good place to draft that. It's not going to happen instantly, but it's a good goal for 2009.

These are the last of my notes from the weekend's Open Hack London event, my notes from various talks are tagged openhacklondon.

Tom Morris, SPARQL and semweb stuff - tech talk at Open Hack London

Tom Morris gave a lightning talk on 'How to use Semantic Web data in your hack' (aka SPARQL and semantic web stuff).

He's since posted his links and queries - excellent links to endpoints you can test queries in.

Semantic web often thought of as long-promised magical elixir, he's here to say it can be used now by showing examples of queries that can be run against semantic web services. He'll demonstrate two different online datasets and one database that can be installed on your own machine.

First - dbpedia - scraped lots of wikipedia, put it into a database. dbpedia isn't like your averge database, you can't draw a UML diagram of wikipedia. It's done in RDF and Linked Data. Can be queried in a language that looks like SQL but isn't. SPARQL - is a w3c standard, they're currently working on SPARQL 2.

Go to dbpedia.org/sparql - submit query as post. [Really nice - I have a thing about APIs and platforms needing a really easy way to get you to 'hello world' and this does it pretty well.]

[Line by line comments on the syntax of the queries might be useful, though they're pretty readable as it is.]

'select thingy, wotsit where [the slightly more complicated stuff]'

Can get back results in xml, also HTML, 'spreadsheet', JSON. Ugly but readable. Typed.

[Trying a query challenge set by others could be fun way to get started learning it.]

One problem - fictional places are in Wikipedia e.g. Liberty City in Grand Theft Auto.

Libris - how library websites should be
[I never used to appreciate how much most library websites suck until I started back at uni and had to use one for more than one query every few years]

Has a query interface through SPARQL

Comment from the audience BBC - now have SPARQL endpoint [as of the day before? Go BBC guy!].

Playing with mulgara, open source java triple store. [mulgara looks like a kinda faceted search/browse thing] Has own query language called TQL which can do more intresting things than SPARQL. Why use it? Schemaless data storage. Is to SQL what dynamic typing is to static typing. [did he mean 'is to sparql'?]

Question from audence: how do you discover what you can query against?
Answer: dbpedia website should list the concepts they have in there. Also some documentation of categories you can look at. [Examples and documentation are so damn important for the update of your API/web service.]

Coming soon [?] SPARUL - update language, SPARQL2: new features

The end!

[These are more (very) rough notes from the weekend's Open Hack London event - please let me know of clarifications, questions, links or comments. My other notes from the event are tagged openhacklondon.

Quick plug: if you're a developer interested in using cultural heritage (museums, libraries, archives, galleries, archaeology, history, science, whatever) data - a bunch of cultural heritage geeks would like to know what's useful for you (more background here). You can comment on the #chAPI wiki, or tweet @miaridge (or @mia_out). Or if you work for a company that works with cultural heritage organisations, you can help us work better with you for better results for our users.]

There were other lightning talks on Pachube (pronounced 'patchbay', about trying to build the internet of things, making an API for gadgets because e.g. connecting hardware to the web is hard for small makers) and Homera (an open source 3d game engine).

Tuesday, 12 May 2009

Mashups made of messages - tech talk at Open Hack London

More (very) rough notes from the weekend's Open Hack London event - please let me know of clarifications, questions, links or comments. You can also check out other posts here tagged openhacklondon.

Mashups made of messages, Matt Biddulph (Dopplr)

Systems architecture on Doppler lets them combine 3rd party systems with their stuff without tying their servers up in knots.

At a rough count, Dopplr uses about 25 third party web APIs.

If you're going to make a web service, site, concentrate on the stuff you're good at. [Use what other people are good at to make yours ace.]

But this also means you're outsourcing and part of your reliability to other people. For each bit of service you add, network latency [is?] putting another bit of risk into your web architecture. Use messaging systems to make server side stuff asynchronous.

'&' is his favourite thing about Linux. Fundamental in Unix that work is divided into packets; each doing the thing it does well. Not even very tightly coupled. Anything that can be run on the command line, stick & on the end, do it in the background. Can forget about things running in the background - don't have to manage the processes, it's not tightly coupled.

Nothing in web apps is simple these days - lots of interconnected bits.

In the physical world, big machines use gearing - having different bits of system run at different speeds. Also things can freewheel then lock in to system again when done.

When building big systems, there's a worry that one machine, one bit it depends on can bring down everything else.

[Slide of a] Diagram of all the bits of the system that don't run because someone has sent an HTTP request - [i.e. background processes]

Flickr is doing less database work up front to make pages load as quickly as possible. They queue other things in the background. e.g. photos load, tags added slightly later. (See post 'Flickr engineers do it offline'.)

Enterprise Integration Patterns (Hohpe et al) is a really good book. Banks have been using messaging for years to manage the problems. Atomic packets of data can be sent on a channel - 'Email for applications'.

Designing - think about what needs to be done now, what can be done in the background? Think of it as part of product design - what has instant effect, what has slower effect? Where can you perform the 'sleight of hand' without people noticing/impacting their user experience?

Example using web services 1: Dopplr and AMEE. What happens when someone asks to see their carbon impact? A request for carbon data goes to Ruby on Rails (memory hungry, not the fastest thing in the world, try to take things off that and process elsewhere). Refresh user screen 'check back soon', send request to message broker (in JSON). Worker process connected to message broker sends request to AMEE. Update database.

Example using web services 2: Flickr pictures on Dopplr page. When you request a trip page, the page loads with all usual stuff and empty div in page with a piece of Javascript on a timer that polls Flickr.

Keeps open connection, a way to push messages to the client while it's waiting to do something.

When processing lots of stuff, worker processes write to memcache as a form of progress bar, but the process is actually disconnected from the webserver so load/risk is outsourced.

'Sites built with glue and string don't automatically scale for free.' You can have many webservers, but the bottleneck might be in the database. Splitting work into message queues is a way of building so things can scale in parallel.

Slide of services, companies that offer messaging stuff. [Did anyone get a photo of that?]

Because of abstraction and with things happening in the background, it's a different flow of control than you might be used to - monitoring is different. You can't just sit there with a single debugger.

[Slide] "If you can't see your changes take effect in a system your understanding of cause and effect breaks down" - not just about it being hard to debug, it's also about user expectations.

I really liked this presentation - it's always good to learn from people who are not only innovating, but are also really solid on performance and reliability as well as the user experience.

[Update: a version of this talk is on the Dopplr blog with slides and notes.]

Saturday, 9 May 2009

Rasmus Lerdorf on Hacking with PHP - tech talk at Open Hack London

Same deal as my first post from today's Open Hack London event - these are (very) rough notes, please let me know of clarifications, questions or comments.

Hacking with PHP, Rasmus Lerdorf

Goal of talk: copy and pastable snippets that just work so you don't have to fight to get things that work [there's not enough of this to help beginners get over that initial hump]. The slides are available at http://talks.php.net/show/openhack and these notes are probably best read as commentary alongside the code examples.

[Since it's a hack day, some] Hack ideas: fix something you use every day; build your own targeted search engine; improve the look of search results; play with semantic web tools to make the web more semantic; tell the world what kind of data you have - if a resume, use hResume or other appropriate microformats/markup; go local - tools for helping your local community; hack for good - make the world a better place.

SearchMonkey and BOSS are blending together a little bit.

What we need to learn
With PHP – enough to handle simple requests; talk to backend datastore; how to parse XML with PHP, how to generate JSON, some basic javasccript, a JavaScript utility library like YUI or jquery.

parsing XML: simpleXML_load_file() - can load entire URL or local file.

Attributes on node show up as array. Namespace attributes call children of node, name namespace as argument.

Now know how to parse XML, can get lots of other stuff.
Context extraction service, Yahoo - doesn't get enough attention. Post all text, gives you back four or five key terms - can then do an image search off them. Or match ads to webpages.

Can use get or post (curl) - usually too much for get.

PHP to JavaScript on initial page load: JSON_encode -> javascript.

Javascript to PHP (and back)
If you can figure out these six lines of code, you can write anything in the world. How every modern web application works.
Server-side php, client-side javascript.

'There's nothing to building web applications, you just have to break everything down into small enough chunks that it all becomes trivial'.

AJAX in 30 seconds.
Inline comments in code would help for people reading it without hearing the talk at the same time.

JavaScript libraries to the rescue
load maps API, create container (div) for the map, then fill it.

Form - on submit call return updateMap(); with new location.

YGeoRSS - if have GeoRSS file... can point to it.

GeoPlanet - assigns a WOE ID to a place. Locations are more than just a lat long - carry way more information. Basically gives you a foreign key. YQL is starting to make the web a giant database. Can make joins across APIs - woeid works as fk.

YQL - 'combines all the APIs on the web into a single API'.

Add a cache - nice to YQL, and also good for demos etc. Copy and paste cache function from his slides - does a local cache on URL. Hashed with md5. Using PHP streams - #defn. Adding a cache speeds up developing when hacking (esp as won't be waiting for the wifi). [This is a pretty damn good tip cos it's really useful and not immediately obvious.]

XPath on URL using PHP's OAuth extension

SearchMonkey - social engineering people into caring about semantic data on the web. For non-geeks, search plug-in mechanism that will spruce up search results page. Encourages people to add semantic data so their search result is as sexy as their competitors - so goal is that people will start adding semantic data.

'If you're doing web stuff, and don't know about microformats, and your resume doesn't have hResume, you're not getting a job with Yahoo.'

Question: how are microformats different to RDFa?
Answer: there are different types of microformats - some very specific ones, eg hResume, hCal. RDFa - adding arbitrary tags to page. even if no specific way to describe your data. But there's a standard set of mark-ups for a resume so can use that. if your data doesn't match anything at microfomats.org then use RDFa or erdf (?).

RDFa, SearchMonkey - tech talks at Open Hack London

While today's Open Hack London event is mostly about the 24-hour hackathon, I signed up just for the Tech Talks because I couldn't afford to miss a whole weekend's study in the fortnight before my exams (stupid exams). I went to the sessions on 'Guardian Data Store and APIs', 'RDFa SearchMonkey', Arduino, 'Hacking with PHP', 'BBC Backstage', Dopplr's 'mashups made of messages' and lightning talks including 'SPARQL and semantic web' stuff you can do now.

I'm putting my rough and ready notes online so that those who couldn't make it can still get some of the benefits. Apologies for any mishearings or mistakes in transcription – leave me a comment with any questions or clarifications.

One of the reasons I was going was to push my thinking about the best ways to provide API-like access to museum information and collections, so my notes will reflect that but I try to generalise where I can. And if you have thoughts on what you'd like cultural heritage institutions to do for developers, let us know! (For background, here's a lightning talk I did at another hack event on happy museums + happy developers = happy punters).

RDFa - now everyone can have an API.
Mark Birkbeck

Going to cover some basic mark-up, and talk about why RDFa is a good thing. [The slides would be useful for the syntax examples, I'll update if they go online.]

RDFa is a new syntax from W3C - a way of embedding metadata (RDF) in HTML documents using attributes.

e.g. <span property="dc:title"> - value of property is the text inside the span.

Because it's inline you don't need to point to another document to provide source of metadata and presentation HTML.

One big advance is that can provide metadata for other items e.g. images, so you can e.g. attach licence info to the image rather than page it's in – e.g. <img src="" rel="licence" resource="[creative commons licence]">

Putting RDFa into web pages means you've now got a feed (the web page is the RSS feed), and a simple static web page can become an API that can be consumed in the same way as stuff from a big expensive system. 'Growing adoption'.

Government department Central Office of Information [?] is quite big on RDFa, have a number of projects with it. [I'd come across the UK Civil Service Job Service API while looking for examples for work presentations on APIs.]

RDFa allows for flexible publishing options. If you're already publishing HTML, you can add RDFa mark-up then get flexible publishing models - different departments can keep publishing data in their own way, a central website can go and request from each of them and create its own database of e.g. jobs. Decentralised way of approaching data distribution.

Can be consumed by: smarter browsers; client-side AJAX, other servers such as SearchMonkey.

He's interested where browsers can do something with it - either enhanced browsers that could e.g. store contact info in a page into your address book; or develop JavaScript libraries that can parse page and do something with it. [screen shot of jobs data in search monkey with enhanced search results]

RDFa might be going into Drupal core.

Example of putting isbn in RDFa in page, then a parser can go through the page, pull out the triples [some explanation of them as mini db?], pull back more info about the book from other APIs e.g. Amazon - full title, thumbnail of cover. e.g. pipes.

Example of FOAF - twitter account marked up in page, can pull in tweets. Could presumably pull in newer services as more things were added, without having to re-mark-up all the pages.

Example of chemist writing a blog who mentions a chemical compound in blog post, a processor can go off and retrieve more info - e.g. add icon for mouseover info - image of molecule, or link to more info.

Next plan is to link with BOSS. Can get back RDFa from search results - augment search results with RDFa from the original page.

Search Monkey (what it is and what you can do with it)
Neil Crosby (European frontend architect for search at Yahoo).

SearchMonkey is (one of) Yahoo's open search platforms (along with BOSS). Uses structured data to enhance search results. You get to change stuff on Yahoo search results page.

SearchMonkey lets you: style results for certain URL patterns; brand those results; make the results more useful for users.

[examples of sites that have done it to see how their results look in Yahoo? I thought he mentioned IMDb but it doesn't look any different - a film search that returns a wikipedia result, OTOH, does.]

Make life better for users - not just what Yahoo thinks results should be, you can say 'actually this is the important info on the page'

Three ways to do it [to change the SERP [search engine results page]: mark up data in a way that Yahoo knows about - 'just structure your data nicely'. e.g. video mark-up; enhance a result directly; make an infobar.

Infobar - doesn't change result see immediately on the page, but it opens on the page. e.g. of auto-enhanced result- playcrafter. Link to developer start page - how to mark it up, with examples, and what it all means.

User-enhanced result - Facebook profile pages are marked up with microformats - can add as friend, poke, send message, view friends, etc from the search results page. Can change the title and abstract, add image, favicon, quicklinks, key/value pairs. Create at [link I can't see but is on slides] Displayed in screen, you fill it out on a template.

Infobar - dropdown in grey bar under results. Can do a lot more, as it's hidden in the infobar and doesn't have to worry people.

Data from: microformats, RDF, XSLT, Yahoo's index, and soon, top tags from delicious.

If no machine data, can write an XSLT. 'isn't that hard'. Lots of documentation on the web.

Examples of things that have been made - a tool that exposes all the metadata known for a page. URL on slide. can install on Yahoo search page, add it in. Use location data to make a map - any page on web with metadata about locations on it - map monkey. Get qype results for anything you search for.

There's a mailing list (people willing and wanting to answer questions) and a tutorial.

Questions

Question: do you need to use a special doctype [for RDFa]?
Answer: added to spec that 'you should use this doctype' but the spec allows for RDFa to be used in situations when can't change doctype e.g. RDFa embedded in blogger blogpost. Most parsers walk the DOM rather than relying on the doctype.

Jim O'D - excited that SearchMonkey supports XSLT - if have website with correctly marked up tables, could expose those as key/value pairs?
Answer: yes. XSLT fantastic tool for when don't have data marked up - can still get to it.

Frankie - question I couldn't hear. About info out to users?
Answer: if you've built a monkey, up to you to tell people about it for the moment. Some monkeys are auto-on e.g. Facebook, wikipedia... possibly in future, if developed a monkey for a site you own, might be able to turn it auto-on in the results for all users... not sure yet if they'll do it or not.
Frankie: plan that people get monkeys they want, or go through gallery?
Answer: would be fantastic if could work out what people are using them for and suggest ones appropriate to people doing particular kinds of searches, rather than having to go to a gallery.