Showing posts with label experimental. Show all posts
Showing posts with label experimental. Show all posts

Monday, 12 November 2012

Reflections on teaching Neatline

I've called this post 'Reflections on teaching Neatline' but I could also have called it 'when new digital humanists meet new software'. Or perhaps even 'growing pains in the digital humanities?'.

A few months ago, Anouk Lang at the University of Strathclyde asked me to lead a workshop on Neatline, software from the Scholar's Lab that plots 'archives, objects, and concepts in space and time'. It's a really exciting project, designed especially for humanists - the interfaces and processes are designed to express complexity and nuance through handcrafted exhibits that link historical materials, maps and timelines.

The workshop was on Thursday, and looking at the evaluation forms, most people found it useful but a few really struggled and teaching it was also slightly tough going. I've been thinking a lot about the possible reasons for that and I'm sharing them both as a request for others to share their experiences in similar circumstances and also in the hope that they'll help others.

The basic outline of the workshop was an intros round (who I am, who they are and what they want to learn); information on what Neatline is and what it can do; time to explore Neatline and explore what the software can and can't do (e.g. login, follow the steps at neatline.org/plugins/neatline to create an item based on a series of correspondence Anouk had been working on, deciding whether you want to transcribe or describe the letter, tweaking its appearance or linking it to other items); and a short period for reflection and discussion (e.g. 'What kinds of interpretive decisions did you find yourself making? What delighted you? What frustrated you?') to finish. If you're curious, you can follow along with my slides and notes or try out the Neatline sandbox site.

The first half was fine but some people really struggled with the hands-on section. Some of it was to do with the software itself - as a workshop, it was a brilliant usability test of the admin interfaces of the software for audiences outside the original set of users. Neatline was only launched in July this year and isn't even in version 2 yet so it's entirely understandable that it appears to have a few functional or UX bugs. The documentation isn't integrated into the interface yet (and sometimes lacks information that is probably part of the shared tacit knowledge of people working on the project) but they have a very comprehensive page about working with Neatline items. Overall, the process of handcrafting timelines and maps for a Neatline exhibit is still closer to 'first, catch your rabbit' than making a batch of ready-mix cupcakes. Neatline is also designed for a particular view of the world, and as it's built on top of other software (Omeka) with another very particular view of the world (and hello, Dublin Core), there's a strong underlying mental model that informs the processes for creating content that is foreign to many of its potential users, including some at the workshop.

But it was also partly because I set the bar too high for the exercises and didn't provide enough structure for some of the group. If I'd designed it so they created a simple Neatline item by closely following detailed instructions (as I have done for other, more consciously tech-for-beginners workshops), at least everyone would have achieved a nice quick win and have something they could admire on the screen. From there some could have tried customising the appearance of their items in small ways, and the more adventurous could have tried a few of the potential ways to present the sample correspondence they were working with to explore the effects of their digitisation decisions. An even more pragmatic but potentially divisive solution might have been to start with the background and demonstration as I did, but then do the hands-on activity with a smaller group of people who were up for exploring uncharted waters. On a purely practical level, I also should have uploaded the images of the letters used in the exercise to my own host so that they didn't have to faff with Dropbox and Omeka records to get an online version of the image to use in Neatline.

And finally it was also because the group had really mixed ICT skills. Most were fine (bar the occasional bug), but some were not. It's always hard teaching technical subjects when participants have varying levels of skill and aptitude, but when does it go beyond aptitude into your attitude about being pushed out of your comfort zone? I'd warned everyone at the start that it was new software, but if you haven't experienced beta software before I guess you don't have the context for understanding what that actually means.

I should make it clear here that I think the participants' achievements outshine any shortcomings - Neatline is a great tool for people working with messy humanities data who want to go beyond plonking markers on Google Maps, and I think everyone got that, and most people enjoyed the chance to play with Neatline.

But more generally, I also wonder if it has to do with changing demographics in the digital humanities - increasingly, not everyone interested in DH is an early, or even a late adopter, and someone interested in DH for the funding possibilities and cool factor might not naturally enjoy unstructured exploration of new software, or be intrigued by trying out different combinations of content and functionality just 'to see what happens'.

Practically, more information for people thinking of attending would be useful - 'if you know x already, you'll be fine; if you know y already, you'll be bored' would be useful in future. Describing an event as 'if you like trying new software, this is for you' would probably help, but it looks like the digital humanities might also now be attracting people who don't particularly like working things out as they go along - are they to be excluded? If using software like this is the onboarding experience for people new to the digital humanities, they're not getting the best first impression, but how do you balance the need for fast-moving innovative work-in-progress to be a bit hacky and untidy around the edges with the desires of a wider group of digital humanities-curious scholars? Is it ok to say 'here be dragons, enter at your own risk'?

Friday, 1 June 2012

Museums and the audience comments paradox

I was at the Imperial War Museum for an advisory board meeting for the Social Interpretation project recently, and had a chance to reflect on my experiences with previous audience participation projects.  As Claire Ross summarised it, the Social Interpretation project is asking: does applying social media models to collections successfully increase engagement and reach?  And what forms of moderation work in that environment - can the audience be trusted to behave appropriately?

One topic for discussion yesterday was whether the museum should do some 'gardening' on the comments.  Participation rates are relatively high but some of the comments are nonsense ('asdf'), repetitive (thousands of variants of 'Cool' or 'sad') or off-topic ('I like the museum') - a pattern probably common to many museum 'have your say' kiosks.  Gardening could involve 'pruning' out comments that were not directly relevant to the question asked in the interactive, or finding ways to surface the interesting comments.  While there are models available in other sectors (e.g. newspapers), I'm excited by the possibility that the Social Interpretation project might have a chance to address this issue for museums.

A big design challenge for high-traffic 'have your say' interactives is providing a quality experience for the audience who is reading comments - they shouldn't have to wade through screens of repeated, vacuous or rude comments to find the gems - while appropriately respecting the contribution and personal engagement of the person who left the comment.

In the spirit of 'have your say', what do you think the solution might be?  What have you tried (successfully or not) in your own projects, or seen working well elsewhere?

Update: the Social Interpretation have posted I iz in ur xhibition trolling ur comments:
"One of the most discussed issues was about what we have termed ‘gardening comments’ but to put it bluntly it’s more a case of should we be ‘curating the visitor voice’ in order to improve the visitor experience? It’s a difficult question to deal with... 
We are at the stage where we really do want to respect the commenter, but also want to give other readers a high value experience. It’s a question of how we do that, and will it significantly change the project?"

If you found this post, you might also be interested in Notes from 'The Shape of Things: New and emerging technology-enabled models of participation through VGC'.

Update, March 2014: I've just been reading a journal article on 'Normative Influences on Thoughtful Online Participation'. The authors set out to test this hypothesis:
'Individuals exposed to highly thoughtful behavior from others will be more thoughtful in their own online comment contributions than individuals exposed to behavior exhibiting a low degree of thoughtfulness.' 
Thoughtful comments were defined by the number of words, how many seconds it took to write them, and how much of the content was relevant to the issue discussed in the original post. And the results? 'We found significant effects of social norm on all three measures related to participants’ commenting behavior. Relative to the low thoughtfulness condition, participants in the high thoughtfulness condition contributed longer comments, spent more time writing them, and presented more issue-relevant thoughts.' To me, this suggests that it's worth finding ways to highlight the more thoughtful comments (and keeping pulling out those 'asdf' weeds) in an interactive as this may encourage other thoughtful comments in turn.

Reference: Sukumaran, Abhay, Stephanie Vezich, Melanie McHugh, and Clifford Nass. “Normative Influences on Thoughtful Online Participation.” In Proceedings of the 2011 Annual Conference on Human Factors in Computing Systems, 3401–10. Vancouver, BC, Canada: ACM, 2011. http://dl.acm.org/citation.cfm?id=1979450.

Sunday, 13 September 2009

Let's push things forward - V&A and British Library beta collections search

The V&A and the British Library have both recently released beta sites for their collections searches.  I'd mentioned the V&A's beta collections search in passing elsewhere, but basically it's great to see such a nicely designed interface - it's already a delight to use and has a simplicity that usually only comes from lots of hard work - and I love that the team were able to publish it as a beta.  Congratulations to all involved!

(I'm thinking about faceted browsing for the Science Museum collections, and it's interesting to see which fields the V&A have included in the 'Explore related objects' panel (example).  I'd be interested to see any usability research on whether users prefer 'inline' links to explore related objects (e.g. in the 'tombstone data' bit to the right of the image) or for the links to appear in a distinct area, as on this site. )

I'm not sure how long it's been live, but the British Library beta catalogue search features a useful 'Refine My Results' panel on the right-hand side of the search results page.  

There's also a 'workspace', where items and queries can be saved and managed.  I think there's a unique purpose for users of the BL search that most sites with 'save your items' functions don't have - you can request items directly from your workspace in advance for delivery when next in the library.  My friendly local British Library regular says the ability to save searches between sessions is immensely useful.  You can also export to delicious, Connotea, RefWorks or EndNote, so your data is transportable, though unfortunately when I tested my notes on an item weren't also exported.  I don't have a BL login so I haven't been able to play with their tagging system.

They've included a link to a survey, which is a useful way to get feedback from their users.

Both beta sites are already useful, and I look forward to seeing how they develop.

Monday, 23 March 2009

Get thee to a wiki - the great API challenge in action

Help us work on an informal, lightweight way of devising shared data, API standards for museum and cultural heritage organisations - museum-api.pbwiki.com is open for business.

You could provide examples of APIs you've used or produced, share your experience as a consumer of web services, tell us about your collections.

Commenting on other people's queries and content is an easy way to get started.  I'd particularly love to hear from curators and collections managers - we should be working together to enable greater access to collections.  If you check it out and none of it makes any sense - be brave and say so!  We should be able to explain what we're doing clearly, or we're not doing it right.

Some background: as announced on the nascent museumdev blog, the Science Museum is looking at releasing an API soon - it'll be project-specific to start with, but we're creating it with the intention of using that as an iterative testing and learning process to design an API for wider use. We could re-invent the wheel, but we'd rather make it easy for people to use what they've learnt using other APIs and other museum collections - the easiest way to do that is to work with other museums and developers. The Science Museum's initial public-facing collections API will be used for a 'mashup competition' based on object metadata from our 'cosmos and culture' gallery.

Speaking of museumdev, I started it as somewhere where I could ask questions, point people to discussions, a home for collections of links and stuff in development.  It's also got random technical bits like 'Tip of the Day: saving web.config as Unicode' because I figure I might as well share my mistakes^H^H^H^H^H^H^H^H learning experiences in the hope that someone, somewhere, benefits.

Sunday, 22 February 2009

Happy developers + happy museums = happy punters (my JISC dev8D talk)

This is a rough transcript of my lightning talk 'Happy developers, happy museums' at JISC's dev8D 'developer happiness' days last week. The slides are downloadable or embedded below. The reason I'm posting this is because I'd still love to hear comments, ideas, suggestions, particularly from developers outside the museum sector - there's a contact form on my website, or leave a comment here.

"In this talk I want to show you where museums are in terms of data and hear from you on how we can be more useful.

If you're interested in updates I use my blog to [crap on a bit, ahem] talk about development at work, and also to call for comment on various ideas and prototypes. I'm interested in making the architecture and development process transparent, in being responsive to not only traditional museum visitors as end users, but also to developers. If you think of APIs as a UI for developers, we want ours to be both usable and useful.

I really like museums, I've worked in three museums (or families of museums) now over ten years. I think they can do really good things. Museums should be about delight, serendipity and answers that provoke more questions.

A recent book, 'How does one become a scientist? : survey on the birth of a Vocation' states that '60% of scientists over 30 and 40% of scientists under 30 note claim, without prompting, that the Palais de la Découverte [a science museum in Paris] triggered their vocation'.

Museums can really have an impact on how people think about the world, how they think about the possibilities of their lives. I think museums also have a big responsibility - we should be curating collections for current and future audiences, but also trying to provide access to the collections that aren't on display. We should be committed to accessibility, transparency, curation, respecting and enabling expertise.

So today I'm here because we want to share our stuff - we are already - but we want to share better.

We do a lot of audience research and know a lot about some of our users, including our specialist users, but we don't know so much about how people might use our data, it's a relatively new thing for us. We're used to saying 'here are objects in a case, interpretation in label', we're not used to saying 'here's unmediated access, access through the back door'.

Some of the challenges for museums: technology isn't that much of a challenge for us on the whole, except that there are pockets of excellence, people doing amazing things on small budgets with limited resources, but there are also a lot of old-fashioned monolithic project designs with big overheads that take a long time to deliver. Lots of people mean well but don't know what's possible - I want to spread the news about lightweight, more manageable and responsive ways of developing things that make sense and deliver results.

We have a lot of data, but a lot of it's crap. Some of what we have is wrong. Some of it was written 100 years ago, so it doesn't match how we'd describe things now.

We face big institutional challenges. Some curators - (though it does depend on the museum) - fear loss of control, fear intellectual vandalism, that mistakes in user-generated content published on museum sites will cause people to lose trust in museums. We have fears of getting the IT wrong (because for a while we did). Funding and metrics are a big issue - we are paid by how many people come through our door or come to our websites. If we're doing a mashup, how do we measure the usage of that? Are we going to cost our organisations money if we can't measure visits and charge back to the government? [This is particularly an issue for free museums in the UK, an interesting by-product of funding structures.]

Copyright is a huge issue. We might not even own an object that appears in our collections, we might not own the rights to the image of our object, or to the reproductions of an image. We might not have asked for copyright clearance at the time when an object was donated, and the cost of tracing it might be too high, so we can't use that object online. Until we come up with a reliable model that reduces the risk to an institution of saying 'copyright unknown', we're stuck.

The following are some ways I can think of for dealing with these challenges...
Limited resources - we can't build an interface to meet every need for every user, but we can provide the content that they'd use. Some of the semantic web talks here have discussed a 'thin layer' of application over data, and that's kind of where we want to go as well.

Real examples to reduce institutional fear and to provide real examples of working agile projects. [I didn't mean strictly 'agile' methodology but generally projects that deliver early and often and can respond to the changing technical and social environment]

Finding ways for the sector to reward intelligent failure. Some museums will never ever admit to making a mistake. I've heard over the past few days that universities can be the same. Projects that are hyped up suddenly aren't mentioned, and presumably it's failed, but no-one [from the project] ever talks about why so we don't learn from those mistakes. 'Fail faster, succeed sooner'.
I'd like to hear suggestions from you on how we could deal with those challenges.

What are museums known for? Big buildings, full of stuff; experts; we make visitors come to us; we're known for being fun; or for being boring.

Museum websites traditionally appear to be about where we are, when we're open, what's on, is there a cafe on site. Which is useful, but we can do a lot more.

Traditionally we've done pretty exhibition microsites, which are nice - they provide an experience of the exhibition before or after your visit. They're quite marketing-led, they don't necessarily provide an equivalent experience and they don't really let you engage with the content beyond the fact that you're viewing it.

We're doing lots of collections online projects, some of these have ended up being silos - sometimes to the extent if we want to get data out of them, we have to screen-scrape our own data. These sites often aren't as pretty, they don't always have the same design and usability budgets (if any).

I think we should stick to what we're really good at - understanding the data (collections), understanding how to mediate it, how to interpret it, how to select things that are appropriate for publication, and maybe open it up to other people to do the shiny pretty things. [Sounds almost like I'm advocating doing myself out of a job!]

So we have lots of objects, images, lots of metadata; our collections databases also include people, events, dates, places, businesses and organisations, lots of qualified information around things like dates, they're not necessarily simple fields but that means they can convey a lot more meaning. I've included that because people don't always realise we have information beyond objects and object metadata. This slide [11 below] is an example of one of the challenges - this box of objects might not be catalogued as individual instruments, it might just be catalogued as a 'box of stuff', which doesn't help you find the interesting objects in the box. Lots of good stuff is hidden in this way.

We're slowly getting there. We're opening up access. We're using APIs internally to share data between gallery interactives and the web, we're releasing them as data points, we're using them to provide direct access to collections. At the moment it still tends to be quite mediated access, so you're getting a lot of interpretation and a fewer number of objects because of the resources required to create really nice records and the information around them.

'Read access' is relatively easy, 'write access' is harder because that's when we hit those institutional issues around authority, authorship. Some curators are vaguely horrified that they might have to listen to what the public have to say and actually take some of it back into their collections databases. But they also have to understand that they can't know everything about their collections, and there are some specialist users who will know everything there is to know about a particular widget on a particular kind of train. We'd like to capture that knowledge. [London Transport Museum have had a good go at that.]

Some random URLs of cool stuff happening in museums [http://dashboard.imamuseum.org/, http://www.powerhousemuseum.com/collection/database/menu.php, http://www.brooklynmuseum.org/opencollection/collections/, http://objectwiki.sciencemuseum.org.uk/] - it's still very much in small pockets, it's still difficult for museum staff to convince people to take what seems like a leap of faith and try these non-traditional things out.

We're taking our content to where people hang out. We're exploring things like Flickr Commons, asking people to tag and comment. Some museums have been updating collections records with information added by the public as a result. People are geo-tagging photos for us, which means you can do 'then and now' mashups without a big metadata enhancement budget.

I'd like to see an end to silos. We are kinda getting there but there's not a serious commitment to the idea that we need to let things go, that we need to make sure that collections online shareable, that they're interoperable, that they can mesh with other things.

Particularly for an education audience, we want to help researchers help themselves, to help developers help others. What else do we have that people might find useful?

What we can do depends on who you are. I could hope that things like enquiry-based learning, mashups, linked data, semantic web technologies, cross-collections searches, faceted browsing to make complex searches easy would be useful, that the concept of museums as a place where information lives – a happy home for metadata mapped around objects and authority records - are useful for people here but I wouldn't want to put words into your mouths.

There's a lot we can do with the technology, but if we're investing resources we need to make sure that they're useful. I can try things in my own time because it's fun, but if we're going to spend limited resources on interfaces for developers then we need to that it's actually going to help some group of people out there.

The philosophy that I'm working with is 'we've got really cool things, but we can have even cooler things if we can share what we have with everyone else'. "The coolest thing to do with your data will be thought of by someone else". [This quote turns out to be on the event t-shirts, via CRIG!] So that said... any ideas, comments, suggestions?"

And that, thankfully, is where I stopped blathering on. I'll summarise the discussion and post back when I've checked that people are ok with me blogging their comments.

[If the slide show below has a brown face on a black background, it's the right one - slideshare's embed seems to have had a hiccup. If it's not that, try viewing it online directly.]



[My slide images include the Easter Egg museum in Kolomyya, Ukraine and 'Laughter in Odd Places' event at the Museum of London.]

This is a quick dump of some of the text from an interview I did at the event, cos I managed to cover some stuff I didn't quite articulate in my talk:

[On challenges for museums:] We need to change institutional priorities to acknowledge the size of the online audience and the different levels of engagement that are possible with the online experience. Having talked to people here, museums also need to do a bit of a sell job in letting people know that we've changed and we're not just great big imposing buildings full of stuff.

[What are the most exciting developments in the museum sector, online?] For digital collections, going outside the walls of the museum using geo-location to place objects in their original context is amazing. It means you can overlay the streets of the city with past events and lives. Outsourcing curation and negotiating new models of expertise is exciting. Overcoming the fear of the digital surrogate as a competitor for museum visits and understanding that everything we do builds audiences, whether digital or physical.

Monday, 26 January 2009

'The strikethrough is the canonical symbol of the Web'

Below is a quote from Wired's Chris Anderson on museum, curatorial authority and the long tail, from a Washington Post report, 'Smithsonian Click-n-Drags Itself Forward' on Smithsonian 2.0 ('A Gathering to Re-Imagine the Smithsonian in the Digital Age').

The quote really covers two issues - making failures and mistakes in public and leaving them there, and training external volunteers and experts to curate parts of collections, because no one curator can be authoritative on everything in their remit: "in exchange for a slight diminution of the credentialed voice for a small number of things, you would get far more for a lot of things".

I suspect this is a false dichotomy - there's a place for both internal and external expertise. The Science Museum object wiki doesn't mean the rest of the collection catalogue and interpretation has no value or relevance. The challenge lies in presenting organisation and user-contributed content in the same interface - can those boundaries be removed? Is it wise to try? And what about taking external content back into the catalogue?

This isn't a new conversation for museum technologists, but it's a conversation I'd love to have with curators. I've never been sure how the technologists who get really excited by the possibilities of sharing content online in various ways can go about working with curators to find the best way of managing it so that the public, the collections and the curators benefit.

Anyway, onto Chris Anderson:
The discovery of the "long tail" principle has implications for museums because it means there is vast room at the bottom for everything. Which means, Anderson said, that curators need to get over themselves. Their influence will never be the same.

"The Web is messy, and in that messiness comes something new and interesting and really rich," he said. "The strikethrough is the canonical symbol of the Web. It says, 'We blew it, but we are leaving that mistake out there. We're not perfect, but we get better over time.' "

If you think that notion gives indigestion to an organization like the Smithsonian -- full of people who have devoted much of their lifetimes to bringing near-perfect luster to some tiny pearl of truth -- you would be correct.

The problem is, "the best curators of any given artifact do not work here, and you do not know them," Anderson told the Smithsonian thought leaders. "Not only that, but you can't find them. They can find you, but you can't find them. The only way to find them is to put stuff out there and let them reveal themselves as being an expert."

Take something like, oh, everything the Smithsonian's got on 1950s Cold War aircraft. Put it out there, Anderson suggested, and say, "If you know something about this, tell us." Focus on the those who sound like they have phenomenal expertise, and invest your time and effort into training these volunteers how to curate. "I'll bet that they would be thrilled, and that they would pay their own money to be given the privilege of seeing this stuff up close. It would be their responsibility to do a good job" in authenticating it and explaining it. "It would be the best free labor that you can imagine."

It didn't go down easily among the thought leaders, who have staked their lives' work on authoritativeness, on avoiding strikethroughs. What about the quality and strength of the knowledge we offer? asked one Smithsonian attendee.

You don't get it, Anderson suggested. "There aren't enough of you. Your skills cannot be invested in enough areas to give that quality."

It's like Wikipedia and the Encyclopedia Britannica, Anderson said. Some Wikipedia entries certainly are not as perfectly polished as the Britannica. But "most of the things I'm interested in are not in the Britannica. In exchange for a slight diminution of the credentialed voice for a small number of things, you would get far more for a lot of things. Something is better than nothing." And right now at the Smithsonian, what you get, he said, is "great" or "nothing."

"Is it our job to be smart and be the best? Or is it our job to share knowledge?" Anderson asked.

Wednesday, 20 August 2008

BBC experimenting with inline links in articles

I noticed the following link when reading a BBC article today:

BBC: We are trialling a new way to allow you to explore background material without leaving the page.
If you turn on inline links, they appear as subtly blue text against the usual grey. Some have icons indicating which site the link relates to (YouTube, Wikipedia), others don't. Links with an icon open the content directly over the article; links without icons open the link in the same window, taking you from the BBC story. Screenshot below:


The 'Read more' links to a page, Story links trial, that says:
For a limited period the BBC News Website is experimenting with clickable links within the body of news stories.

If you click on one of these links, a window will appear containing background material relevant to that word that is highlighted. The links have been carefully chosen by our journalists.

We are doing this trial because we want to see if you enjoy exploring background material presented in this way. It's part of our continuing efforts to provide the best possible experience.

In addition to background material from the BBC News website, we are also displaying content from other sites, including Wikipedia, You Tube and Flickr.
I'd be really interested to know what the results of the trial are, and I hope the BBC share them. I've been thinking about inline links and faceted browsing for collections sites recently, and while the response would presumably vary if the links were only to related content on the same site, it would be useful to know how the two types of links are received.

The story I noticed the link on is also interesting because it shows how content created in a 'social software' way can be (probably wilfully, in this case) misinterpreted:

"Downing Street has been accused of wasting taxpayers' money after making a jokey video in response to a petition for Jeremy Clarkson to be made PM.
...
A Conservative Party spokesman said: "While the British public is having to tighten its belts, the government is spending taxpayers' money on a completely frivolous project.""

Thursday, 17 July 2008

Giant squid dissection via live video

I've been watching the recording of the live stream of the first ever public dissection by Museum scientists of a giant squid.

Congratulations to everyone involved at Museum Victoria, it's a great use of technology and a great approach to openness. The explanations were beautifully clear, and did a great job of contextualising the research, the process and the animal itself.

I love the paparazzi-style photo flashes as they rolled the trolley out onto the main floor.

Wednesday, 9 July 2008

20% time - an experiment (with some results)

A company called Atlassian have been experimenting with allowing their engineers 20% of their time to work on free or non-core projects (a la Google). They said:
You see, while everyone knows about Google's 20% time and we've heard about all the neat products born from it (Google News, GMail etc) - we've found it extremely difficult to get any hard facts about how it actually works in practice.
So they started with a list of questions they wanted to answer through their experiment, and they've been blogging about it at http://blogs.atlassian.com/developer/20_percent_time/. It makes for interesting reading, and it's great to see some real evidence starting to emerge.

Hat tip: Tech-Ed Collisions.

Friday, 4 July 2008

How to build a web application in four days

There's been a bit of buzz around 'How To Build A Web App in Four Days For $10,000'. Not everything is applicable to the kinds of projects I'd be involved in, but I really liked these points:
  • The best boost you can give you or your team is to provide the time to be creative.
  • You'll come back to your current projects with a new perspective and renewed energy.
  • It will push your team to learn new skills.
  • Simplify the site and app as much as possible. Try launching with just 'Home', 'Help' and 'About'.
  • Make sure to build on a great framework
  • Be technologically agnostic. If your developers are saying it should be built in a certain language and framework and they have solid reasons, trust them and move on.
  • Coordinate how your designers and developers are going to work together.
  • Get your 'Creation Environment' setup correctly. [See the original post for details]

"The coolest thing to be done with your data will be thought of by someone else"

I discovered this ace quote, "the coolest thing to be done with your data will be thought of by someone else", on JISC's Common Repository Interfaces Group (CRIG) site, via the The Repository Challenge. The CRIG was created to "help identify problem spaces in the repository landscape and suggest innovative solutions. The CRIG consists of a core group of technical, policy and development staff with repository interface expertise. It encourages anyone to join who is dedicated and passionate about surfacing scholarly content on the web."

Read 'repository or federated search' for 'repository' (or think of a federated search as a pseudo-repository) and 'scholarly' for 'cultural heritage' content, and it sounds like an awful lot of fun.

It's also the sentiment behind the UK Government's Show Us a Better Way, the Mashed Museum days and a whole bunch of similar projects.

Wednesday, 2 July 2008

What would you create with public (UK) information?

Show Us a Better Way want to know, and if your idea is good they might give you £20,000 to develop it to the next level.
Do you think that better use of public information could improve health, education, justice or society at large?

The UK Government wants to hear your ideas for new products that could improve the way public information is communicated.
Importantly, you don't need to be a geek:
You don't have to have any technical knowledge, nor any money, just a good idea, and 5 minutes spare to enter the competition.
And they've made "gigabytes of new or previously invisible public information" available for the project, including health, crime and education data (but no personal information).

Wednesday, 25 June 2008

Lonely Planet launch API

Lonely Planet launched their 'Explore API' and developer network at the BBC Mashed 2008 day. Available content includes 'destination content, including geocoded points of interest reviews, destination profiles, traveller-created "best of" lists and travel photographs' from their image library so as a travel junkie I'm already itching to have a play.

There's more background in this interview with Chris Heilmann and Chris Boden on the Yahoo! Developer blog, 'Lonely Planet starts developer program at mashed08 in London', but I thought it was worth pulling out this quote about the benefit of APIs, particularly as they're an organisation whose business model relies on its reputation and content:

Where do you see the benefit in releasing an API? How do you plan to monetize it or is it a loss-leader for you?

We don't have a funky web app like Twitter or Dopplr at this stage but we do have content - in a sense, that content is our platform. We want to take the Lonely Planet content and community experience onto relevant new platforms and make it accessible to travellers in new ways. We're not going to be able to do all of that on our own so we're looking to tap into external sources of innovation and creativity through open collaboration to help us imagine and execute the next generation of services that might enrich the lives of our community.

In terms of monetization, we'll look to work commercially with those developers who come up with innovations that we believe have the potential to create commercial value.

Monday, 23 June 2008

Quick and light solutions at 'UK Museums on the Web Conference 2008'

These are my notes from session 4, 'Quick and light solutions', of the UK Museums on the Web Conference 2008. In the interests of getting my notes up quickly I'm putting them up pretty much 'as is', so they're still rough around the edges. There are quite a few sections below which need to be updated when the presentations or photos of slides go online. [These notes would have been up a lot sooner if my laptop hadn't finally given up the ghost over the weekend.]

Frankie Roberto, 'The guerrilla approach to aggregating online collections'
He doesn't have slides, he's presenting using Firefox 3. [You can also read Frankie's post about his presentation on his blog.]

His projects came out of last year's mashed museum day, where the lack of re-usable cultural heritage data online was a real issue. Talk in the pub turned to 'the dark side' of obtaining data - screen scraping was one idea. Then the idea of FoI requests came up, and Frankie ended up sending Freedom of Information requests to national museums in any electronic format with some kind of structure.

He's not showing site he presented at Montreal, it should be online soon and he'll release the code.

Frankie demonstrated the Science Museum object wiki.

[I found 'how it works' as focus of the object text on the Science Museum wiki a really interesting way of writing object descriptions, it could work well for other projects.]

He has concerns about big top down projects so he's suggesting five small or niche projects. He asked himself, how do people relate to objects?
1. Lots of people say, "I've got one of these" so: ivegotoneofthose.com - put objects up, people can hit button to say 'I have one of those'. The raw numbers could be interesting.
[I suggested this for Exploring 20th Century London at one point, but with a bit more user-generated content so that people could upload photos of their object at home or stories about how they got it, etc. I suppose ivegotoneofthose.com could be built so that it also lets people add content about their particular thing, then ideally that could be pulled back into and displayed on a museum site like Exploring. Would ivegotoneofthose.com sit on top of a federated collections search or would it have its own object list?]
2. Looking at TheyWorkForYou.com, he suggests: TheyCollectForYou.com - scan acquisition forms, publish feeds of which curators have bought what objects. [Bringing transparency to the acquisition process?]
3. Looking at howstuffworks.com, what about howstuffworked.com?
4. 'what should we collect next?' - opening up discourse on purchasing. Frankie took the quote from Indiana Jones: thatbelongsinamuseum.com - people can nominate things that should be in a museum.
5. pricelessartefact.com - [crowdsourcing object evaluation?] - comparing objects to see which is the most valuable, however 'valuable' is defined.
[Except that possibly opens the museum to further risk of having stuff nicked to order]

Fiona Romeo, 'Different ways of seeing online collections'
I didn't take many detailed notes for this paper, but you can see my notes on a previous presentation at Notes from 'Maritime Memorials, visualised' at MCG's Spring Conference.

Mapping - objects don't make a lot of sense about themselves, but are compelling as part of information about an expedition, or failed expedition.

They'll have new map and timeline content launching next month.

Stamen can share information about how they did their geocoding and stuff.

Giving your data out for creative re-use can be as easy as giving out a CSV file.
You always want to have an API or feed when doing any website.
The National Maritime Museum make any data set they can find without licensing restrictions and put it online for creative re-use.

[Slide on approaches to data enhancement.]
Curation is the best approach but it's time-consuming.

Fiona spoke about her experiments at the mashed museum day - she cut and paste transcript data into IBM's Many Eyes. It shows that really good tools are available, even if you don't have resources to work with a company like Stamen.

Mike Ellis presented a summary of the 'mashed museum' day held the day before.

Questions, wrap up session
Jon - always assume there (should be) an API

[A question I didn't ask but posted on twitter: who do we need to get in the room to make sure all these ideas for new approaches to data, to aggregation and federation, new types of experiences of cultural heritage data, etc, actually go somewhere?]

Paul on fears about putting content online: 'since the state of Florida put pictures of their beaches on their website, no-one goes to the beach anymore'.

Metrics:
Mike: need to go shout at DCMS about the metrics, need to use more meaningful metrics especially as thinking of something like APIs
Jon: watermark metadata... micro-marketing data.
Fiona: send it out with a wrapper. Make it embeddable.

Question from someone from Guernsey Museum about images online: once you've downloaded your nice image its without metadata. George: Flickr like as much data in EXIF as possible. EXIF data isn't permanent but is useful.

Angela Murphy: wrappers are important for curators, as they're more willing to let things go if people can get back to the original source.

Me, referring back to the first session of the day: what were Lee Iverson's issues with the keynote speech? Lee: partly about the role of institution like the BBC in modern space. National broadcaster should set social common ground, be a fundamental part of democratic discussion. It's even more important now because of variety of sources out there, people shutting off or being selective about information sources to cope with information overload. Disparate source mean no middle ground or possibility of discussion. BBC should 'let it go' - send the data out. The metric becomes how widely does it spread, where does it show up? If restricted to non-commercial use then [strangling use/innovation].

The 'net recomender' thing is a flawed metric - you don't recommend something you disagree with, something that is new or difficult knowledge. What gets recommended is a video of a cute 8 year old playing Guitar Hero really well. People avoid things that challenge them.

Fiona - the advantage of the 'net recomender' is it's taking judgement of quality outside originating institution.

Paul asked who wondered why 7 - 8 on scale of 10 is neutral for British people, would have thought it's 5 - 6.

Angela: we should push data to DCMS instead of expecting them to know what they could ask for.

George: it's opportunity to change the way success is measured. Anita Roddick says 'when the community gives you wealth, it's time to give it back'. [Show, don't tell] - what would happen if you were to send a video of people engaging instead of just sending a spreadsheet?

Final round comments
Fiona: personal measure of success - creating culture of innovation, engagement, creating vibrant environment.

Paul: success is getting other people to agree with what we've been talking about [at the mashed museum day and conference] the past two days. [yes yes yes!] A measure of success was how a CEO reacted to discovering videos about their institution on YouTube - he didn't try to shut it down, but asked, 'how we can engage with that'

Ross on 'take home' ideas for the conference
Collections - we conflate many definitions in our discussions - images, records, web pages about collections.

Our tone has changed. Delivery changed - realignment of axis of powers, MLA's Digital portfolio is disappearing, there's a vacuum. Who will fill it? The Collections Trust, National Museum Directors' Conference? Technology's not a problem, it's the cultural, human factors. We need to talk about where the tensions are, we've been papering over the cracks. Institutional relationships.

The language has changed - it was about digitisation, accessibility, funding. Three words today - beauty, poetry, life. We're entering an exciting moment.

What's the role of the Museums Computer Group - how and what can the MCG do?

Thursday, 15 May 2008

Notes from 'Aggregating Museum Data – Use Issues' at MW2008

These are my notes from the session 'Aggregating Museum Data – Use Issues' at Museums and the Web, Montreal, April 2008.

These notes are pretty rough so apologies for any mistakes; I hope they're a bit useful to people, even though it's so late after the event. I've tried to include most of what was covered but it's taken me a while to catch up on some of my notes and recollection is fading. Any comments or corrections are welcome, and the comments in [square brackets] below are me. All the Museums and the Web conference papers and notes I've blogged have been tagged with 'MW2008'.

This session was introduced by David Bearman, and included two papers:
Exploring museum collections online: the quantitative method by Frankie Roberto and Uniting the shanty towns - data combining across multiple institutions by Seb Chan.

David Bearman: the intentionality of the production of data process is interesting i.e. the data Frankie and Seb used wasn't designed for integration.

Frankie Roberto, Exploring museum collections online: the quantitative method (slides)
He didn't give a crap of the quality of the data, it was all about numbers - get as much as possible to see what he could do with it.

The project wasn't entirely authorised or part of his daily routine. It came in part from debates after the museum mash-up day.

Three problems with mashing museum data: getting it, (getting the right) structure, (dealing with) dodgy data

Traditional solutions:
Getting it - APIs
Structure - metadata standards
Dodgy data - hard work (get curators to fix it)

But it doesn't have to be perfect, it just has to be "good enough". Or "assez bon" (and he hopes that translation is good enough).

Options for getting it - screen scrapers, or Freedom of Information (FOI) requests.

FOI request - simple set of fields in machine-readable format.

Structure - some logic in the mapping into simple format.

Dodgy data - go for 'good enough'.

Presenting objects online: existing model - doesn't give you a sense of the archive, the collection, as it's about the individual pages.

So what was he hoping for?
Who, what, where, when, how. ['Why' is the other traditional journalists questions but too difficult in structured information]

And what did he get?
Who: hoping for collection/curator - no data.
What: hoping for 'this is an x'. Instead got categories (based on museum internal structures).
Where: lots of variation - 1496 unique strings. The specificity of terms varies on geographic and historical dimensions.
When: lots of variation
How: hoping for donation/purchase/loan. Got a long list of varied stuff.

[There were lots of bits about whacking the data together that made people around me (and me, at times) wince. But it took me a while to realise it was a collection-level view, not an individual object view - I guess that's just a reflection of how I think about digital collections - so that doesn't matter as much as if you were reading actual object records. And I'm a bit daft cos the clue ('quantitative') was in the title.

A big part of the museum publication process is making crappy date and location and classification data correct, pretty and human-readable, so the variation Frankie found in data isn't surprising. Catalogues are designed for managing collections, not for publication (though might curators also over-state the case because they'd always rather everything was tidied than published in a possible incorrect or messy state?).

It would have been interesting to hear how the chosen fields related to the intended audience, but it might also have been just a reasonable place to start - somewhere 'good enough' - I'm sure Frankie will correct me if I'm wrong.]

It will be on museum-collections.org. Frankie showed some stuff with Google graph APIs.

Prior art - Pitt Rivers Museum - analysis of collections, 'a picture of Englishness'.

Lessons from politics: theyworkforyou for curators.

Issues: visualisations count all objects equally. e.g. lots of coins vs bigger objects. [Probably just as well no natural history collections then. Damn ants!]

Interactions - present user comments/data back to museums?

Whose role is it anyway, to analyse collections data? And what about private collections?

Sebastian Chan, Uniting the shanty towns - data combining across multiple institutions (slides)
[A paraphrase from the introduction: Seb's team are artists who are also nerds (?)]

Paper is about dealing with the reality of mixing data.

Mess is good, but... mess makes smooshing things together hard. Trying to agree on standards takes a long time, you'll never get anything built.

Combination of methods - scraping + trust-o-meter to mediate 'risk' of taking in data from multiple sources.

Semantic web in practice - dbpedia.

Open Calais
- bought out from Clearforest by Reuters. Dynamically generated metadata tags about 'entities' e.g. possible authority records. There are problems with automatically generated data e.g. guesses at people, organisations, whatever might not be right. 'But it's good enough'. Can then build onto it so users can browse by people then link to other sites with more information records about them in other datasets.

[But can museums generally cope with 'good enough'? What does that do to ideas of 'authority'? If it's machine-generated because there's not enough time for a person in the museum to do it, is there enough time for a person in the museum to clean it? OTOH, the Powerhouse model shows you can crowdsource the cleaning of tags so why not entities. And imagine if we could connect Powerhouse objects in Sydney with data about locations or people in London held at the Museum of London - authority versus utility?

Do we need to critically examine and change the environment in which catalogue data is viewed so that the reputation of our curators/finds specialists in some of the more critical (bitchy) or competitive fields isn't affected by this kind of exposure? I know it's a problem in archaeology too.]

They've published an OpenSearch feed as GeoRSS.

Fire eagle, Yahoo beta product. Link it to other data sets so you can see what's near you. [If you can get on the beta.]

I think that was the end, and the next bits were questions and discussion.

David Bearman: regarding linked authority files... if we wait until everything is perfect before getting it out there, then "all curators have to die before we can put anything on the web", "just bloody experiment".

Nate (Walker): is 'good enough' good enough? What about involving museums in creating better and correcting data. [I think, correct me if not]
Seb: no reason why a museum community shouldn't create an OpenCalais equivalent. David: Calais knows what reuters know about data. [So we should get together as a sector, nationally or internationally, or as art, science, history museums, and teach it about museum data.]

David - almost saying 'make the uncertainty an opportunity' in museum data - open it up to the public as you may find the answers. Crowdsource the data quality processes in cataloguing! "we find out more by admitting we know less".

Seb - geo-location is critical to allowing communities to engage with this material.

Frankie - doing a big database dump every few months could be enough of an API.

Location sensitive devices are going to be huge.

Seb - we think of search in a very particular way, but we don't know how people want to search i.e. what they want to search for, how they find stuff. [This is one of the sessions that made me think about faceted browsing.]

"Selling a virtual museum to a director is easier than saying 'put all our stuff there and let people take it'".

Tim Hart (Museum Victoria) - is the data from the public going back into the collection management system? Seb - yep. There's no field in EMu for some of the stuff that OpenCalais has, but the use of it from OpenCalais makes a really good business case for putting it into EMu.

Seb - we need tools to create metadata for us, we don't and won't have resources to do it with humans.

Seb - Commons on Flickr is good experiment in giving stuff away. Freebase - not sure if go to that level.

Overall, this was a great session - lots of ideas for small and large things museums can do with digital collections, and it generated lots of interesting and engaged discussion.

[It's interesting, we opened up the dataset from Çatalhöyük for download so that people could make their own interpretations and/or remix the data, but we never got around to implementing interfaces so people could contribute or upload the knowledge they created back to the project, or how to use the queries they'd run.]

Wednesday, 7 May 2008

Let's help our visitors get lost

In 'Community: From Little Things, Big Things Grow' on ALA, George Oates from Flickr says:

It's easy to get lost on Flickr. You click from here to there, this to that, then suddenly you look up and notice you've lost hours. Allow visitors to cut their own path through the place and they'll curate their own experiences. The idea that every Flickr visitor has an entirely different view of its content is both unsettling, because you can't control it, and liberating, because you’ve given control away. Embrace the idea that the site map might look more like a spider web than a hierarchy. There are natural links in content created by many, many different people. Everyone who uses a site like Flickr has an entirely different picture of it, so the question becomes, what can you do to suggest the next step in the display you design?

I've been thinking about something like this for a while, though the example I've used is Wikipedia. I have friends who've had to ban themselves from Wikipedia because they literally lose hours there after starting with one innocent question, then clicking onto an interesting link, then onto another...

That ability to lose yourself as you click from one interesting thing to another is exactly what I want for our museum sites: our visitor experience should be as seductive and serendipitous as browsing Wikipedia or Flickr.

And hey, if we look at the links visitors are making between our content, we might even learn something new about our content ourselves.

Saturday, 3 May 2008

MultiMimsy database extractions and the possibilities for OAI-based collections repositories

I've uploaded my presentation slides from a talk for the UK MultiMimsy Users group in Docklands last month to MultiMimsy database extractions and the possibilities for OAI-based collections repositories at the Museum of London.

The first part discusses how to get from a set of data in a collections management system to a final published website, looking at the design process and technical considerations. Willoughby's use of Oracle on the back-end means that any ODBC-compliant database can query the underlying database and extract collections data.

The paper then looks at some of the possibilities for the Museum of London's OAI-PMH repository. We've implemented an OAI repository for the People's Network Discover Service (PNDS) for Exploring 20th Century London (which also means we're set to get records into Europeana), but I hope that we can use the repository in lots of other ways, including the possibility of using our repository to serve data for federated searches.

There's currently some discussion internationally in the cultural heritage sector about repositories vs federated search, but I'm not sure it's an either/or choice. The reasons each are used are often to do with political or funding factors instead of the base technology, but either method, or both, could be used internally or externally depending on the requirements of the project and institution.

I can go into more detail about the scripts we use to extract data from MultiMimsy or send sample scripts if people are interested. They might be a good way to get started if you haven't extracted data from MultiMimsy before but they won't generally be directly relevant to your data structres as the use of MultiMimsy can vary so widely between types of museums, collections and projects.

Thursday, 17 April 2008

Calling geeks in the UK with an interest in cultural heritage content/audiences

You might be interested in BathCamp - a bar camp in Bath on a Saturday (with overnight stay) in late August. This is an initial open call so head along to the website (BathCamp) and check it out. Ideally you would have an interest in cultural heritage content, audiences or applications, but we love the idea of getting fresh perspectives from a wide range of people so we don't expect that you would have worked with the cultural heritage sector (museums, galleries, libraries, archives, archaeology) before.

Tuesday, 15 April 2008

How I do documentation: a column of bumph and a column of gold

All programmers hate documentation, right? But I've discovered a way to make it less painful and I'm posting in case it helps anyone else.

The first trick is to start documenting as soon as you start thinking about a project - well before you've written any code. I keep a running document of the work I've done, including the bits I'm about to try, information about links into other databases or applications, issues I need to think about or questions I need to ask someone, rude comments (I know, I look like such a nice girl), references, quick use cases, bits about functions, summary notes from meetings, etc.

Mostly I record by date, blog style. Doing it by date helps me link repository files, paper notes and emails with particular bits of work, which can otherwise be tricky if it's a while since you worked on a project or if you have lots of projects on the go. It's also handy if you need to record the time spent on different projects.

I just did it like this for a while, and it was ok, but I learnt the hard way that it takes a while to sort through it if I needed to send someone else some documentation. Then I made a conscious decision to separate the random musings from the decisions and notes on the productive bits of code.

So now my document has two columns. This first column is all the bumph described above - the stuff I'd need if I wanted to retrace my steps or remind myself why I ended up doing things a certain way. The second column records key decisions or final solutions. This is your column of gold.

This way I can quickly run down the items in the second column, organise it by area instead of by date and come up with some good documentation without much effort. And if I ever want to write up the whole project, I've got a record of the whole process in the column of bumph.

You could add a third column to record outstanding tasks or questions. I tend to mark these up with colour and un-colour them when they're done. It just depends how you like to work.

It's amazingly simple, but it works. I hope it might be useful for you too. Or if you have any better suggestions (or a better title for this post), I'd love to hear them.

Friday, 7 March 2008

Move your FAQ to Wikipedia?

Mal Booth from the Australian War Memorial (AWM) makes the fascinating suggestion: they should move their entire Encyclopaedia to Wikipedia. Their encyclopaedia seems to function as a fully researched and referenced FAQ with content creation driven by public enquiries, and would probably sit well in Wikipedia.

In Wikipedia and "produsers", Mal says:
"Putting the content up on Wikipedia.org gives it MUCH wider exposure than our website ever can and it therefore has the potential to bring new users to our website that may not even know we exist (via links in to our own web content). With a wikipedia.org user account, we can maintain an appropriate amount of control over the content (more than we have at present over wikipedia content that started as ours, already put up there by others).

Another point is that putting it up on Wikipedia allows us to engage the assistance of various volunteers who'd like to help us, but don't live locally."
He also presents some good suggestions from their web developer, Adam: they should understand and participate in the Wikipedia community, and identify themselves as AWM professionals before importing content. I think they've taken the first step by assessing the suitability of their content for Wikipedia.

It's also an interesting example of an organisation that is willing to 'let go' of their content and allow it to be used and edited outside their institution. Mal's blog is a real find (and I'm not just saying that because it has 'Melbin' (Melbourne) in the title), and I'll be following the progress of their project with interest.

I wonder how issues of trust and authority will play out on their entries: by linking to the relevant Wikipedia entries, the AWM is giving those entries a level of authority they might not otherwise have. They're also placing a great deal of trust in Wikipedia authors.

Mal links to a post by Alex Bruns, Beyond Public Service Broadcasting: Produsage at the ABC and summarises the four preconditions for good user-generated content:
  • the replacement of a hierarchy with a more open participatory structure;
  • recognising the power of the COMMUNITY to distinguish between constructive and destructive contributions;
  • allowing for random (granular, simple) acts of participation (like ratings); and
  • the development of shared rather than owned content that is able to be re-used, re-mixed or mashed up.
Adam's post lists key principles that anyone "looking to develop successful and sustainable participatory media environments" should take into account. These points are defined and expanded on in the original post, which is well worth reading:
  1. Open Participation, Communal Evaluation
  2. Fluid Heterarchy, Ad Hoc Meritocracy
  3. Unfinished Artefacts, Continuing Process
  4. Common Property, Individual Rewards