Showing posts with label tagging. Show all posts
Showing posts with label tagging. Show all posts

Tuesday, 9 September 2014

Helping us fly? Machine learning and crowdsourcing

Moon Machine by Bernard Brussel-Smith via Serendip-o-matic
Over the past few years we've seen an increasing number of projects that take the phrase 'human-computer interaction' literally (or perhaps turning HCI into human-computer integration), organising tasks done by people and by computers into a unified system. One of the most obvious benefits of crowdsourcing on digital platforms has been the ability to coordinate the distribution and validation of tasks, but now data classified by people through crowdsourcing is being fed into computers to improve machine learning so that computers can learn to recognise images almost as well as we do. I've outlined a few projects putting this approach to work below. Of course, this creates new challenges for the future - what do cultural heritage crowdsourcing projects do when all the fun tasks like image tagging and text transcription can be done by computers? After all, Fast Company reports 'at least one Zooniverse project, Galaxy Zoo Supernova, has already automated itself out of existence'. More positively, assuming we can find compelling reasons for people to spend time with cultural heritage collections, how does machine learning and task coordination free us to fly further?

The Public Catalogue Foundation has taken tags created through Your Paintings Tagger and turned them over to computers. As they explain, the results are impressive. The art of computer image recognition: 'Using the 3.5 million or so tags provided by taggers, the research team at Oxford 'educated' image-recognition software to recognise the top tagged terms. Professor Zisserman explains this is a three stage process. Firstly, gather all paintings tagged by taggers with a particular subject (e.g. ‘horse’). Secondly, use feature extraction processes to build an ‘object model’ of a horse (a set of characteristics a painting might have that would indicate that a horse is present). Thirdly, run this algorithm over the Your Paintings database and rank paintings according to how closely they match this model.'

The BBC World Service archive ‘used an open-source speech recognition toolkit to listen to every programme and convert it to text’, extracted keywords or tags from the transcripts then got people to check the correctness of the data created: ‘As well as listening to programmes in the archive, users can view the automatic tags and vote on whether they’re correct or incorrect or add completely new tags. They can also edit programme titles and synopses, select appropriate images and name the voices heard’. From Algorithms and Crowd-Sourcing for Digital Archives by Tristan Ferne. See also What we learnt by crowdsourcing the World Service archive by Yves Raimond, Michael Smethurst, Tristan Ferne on 15 September 2014: 'we believe we have shown that a combination of automated tagging algorithms and crowdsourcing can be used to publish a large archive like this quickly and efficiently'.

And of course the Zooniverse is working on this. From their Milky Way project blog, New MWP paper outlines the powerful synergy between citizens scientists, professional scientists, and machine learning: '...a wonderful synergy that can exist between citizen scientists, professional scientists, and machine learning. The example outlined with the Milky Way Project is that citizens can identify patterns that machines cannot detect without training, machine learning algorithms can use citizen science projects as input training sets, creating amazing new opportunities to speed-up the pace of discovery. A hybrid model of machine learning combined with crowdsourced training data from citizen scientists can not only classify large quantities of data, but also address the weakness of each approach if deployed alone.'

The CUbRIK project combines 'machine, human and social computation for multimedia search'. You can try it out at their technical demonstrator, HistoGraph and try to 'collaboratively identify missing information' about historic photographs. [Added October 2014]

If you're interested in the theory, an early discussion of human input into machine learning is in Quinn and Bederson's 2011 Human Computation: A Survey and Taxonomy of a Growing Field. More recently, the SOCIAM: The Theory and Practice of Social Machines project is looking at 'a new kind of emergent, collective problem solving, in which we see (i) problems solved by very large scale human participation via the Web, (ii) access to, or the ability to generate, large amounts of relevant data using open data standards, (iii) confidence in the quality of the data and (iv) intuitive interfaces', including 'citizen science social machines'. If you're really keen, you can get a sense of the state of the field from various conference papers, including ICML ’13 Workshop: Machine Learning Meets Crowdsourcing and ICML ’14 Workshop: Crowdsourcing and Human Computing. There's also a mega-list of academic crowdsourcing conferences and workshops, though it doesn't include much on the tiny corner of the world that is crowdsourcing in cultural heritage.

NB: this post is a bit of a marker so I've somewhere to put thoughts on machine learning and human-computer integration as I finish my thesis; I'll update this post as I collect more references. Do you know of examples I've missed, or implications we should consider? Comment here or on twitter to start the conversation... 

Wednesday, 29 July 2009

Pre-tagging content for sharing on Twitter

The Age newspaper has implemented an interesting social media widget on their article pages. Their 'Join the conversation' widget shows how many people are reading the same article, links to discussion of the page on Twitter, allows the reader to easily add their own comment to Twitter, lists other articles that people who read this article read, and where appropriate, adds a 'Related Coverage' section above the other standard nav links.

I've included a screenshot because I assume attention for each article is fairly transitory so there may not be other readers or twitter discussions on the sample article by the time you read it:

I'm particularly interested in their approach to Twitter. They've used automatically generated hash tags to group together discussion of each article on Twitter. For example, in this article, 'First climate refugees start move to new island home' the hash tag is '#fd-e06x'. If you use their 'Comment on Twitter' link it automatically sends a status to the Twitter site (if you're logged in) with the article URL and '#fd-e06x'.

The 'Read tweets' link takes you to the Twitter search page with the pre-populated search term '#fd-e06x'. Of course, the search won't show any discussion about the issues in the article or about the article itself that haven't used their hash tag.

The system also seems to generate a new hash tag for each article, even those that are updates on previous breaking news stories. So these two articles (Woman hit by train fights for life, Woman fights for life after train station accident) about the same incident have different hash tags - perhaps this can be fixed in a later iteration.

I wonder if it would be possible to harvest possible topical tags from other tweets about the article (checking all the various URL shortening services) to suggest more human-friendly tags related to an article? It's probably not worth it for The Age, but it might be for content with a longer life. Or would an organisation be at risk of appearing to endorse those labels?

Interestingly, the pages don't mention the hash tags, so the process is invisible to the user. Would explaining it lead to greater uptake? I use a Twitter client (partly because I can easily shorten long links, while The Age doesn't pre-shorten their links) so would take the URL directly from the location bar, missing out their hash tag.

Selfishly, because I'm thinking about it for work, I'd love to know how they generate their 'other people who read this' and 'Related Coverage' links. I assume the latter is manually generated, either as direct links or based on article or section metadata.

Also selfishly, I'd like to know their motives, and whether they have any metrics for the success of the project.

Saturday, 13 December 2008

QR tags edging towards mainstream?

A London-based 'tech PR' blog post said this week:
When Kelly Brooks starts appearing in ads featuring QR codes you know that the 2D dot matrix bar code technology is close to a tipping point. Brooks features in a Pepsi campaign that has gone live this week and images of her clutching a QR code have featured in most of the tabloids.
Source: QR codes and the Kelly Brooks Pepsi campaign, hat tip for link: Heleana Quartey.

See also p8tch.com who says 'think of it as a TinyURL you can wear' and emmacott.com who say 'wear your profile'. There's even a Facebook 'add to friends' QR app and a Google Charts QR Code API.

It's interesting timing, as QR codes were discussed in a MCG thread on 'Putting web addresses on interpretation' that was in essence about linking from the offline physical world and online content.

While they're not mainstream enough to be a viable solution yet, we could be getting close to the tipping point where QR tags might become a viable way of bookmarking real world objects and locations. QR tags also provide a way of linking locations to online content without the requirements for a location-aware device.

Friday, 24 October 2008

Finding problems for QR tags to solve

QR tags (square or 2D barcodes that can hold up to 4,296 characters) are famously 'big in Japan'. Outside of Japan they've often seemed a solution in search of a problem, but we're getting closer to recognising the situations where they could be useful.

There's a great idea in this blog post, Video Print:
By placing something like a QR code in the margin text at the point you want the reader to watch the video, you can provide an easy way of grabbing the video URL, and let the reader use a device that's likely to be at hand to view the video with...
I would use this a lot myself - my laptop usually lives on my desk, but that's not where I tend to read print media, so in the past I've ripped URLs out of articles or taken a photo on my phone to remind myself to look at them later, but I never get around to it. But since I always have my phone with me I'd happily snap a QR code (the Nokia barcode software is usually hidden a few menus down, but it's worth digging out because it works incredibly well and makes a cool noise when it snaps onto a tag) and use the home wifi connection to view a video or an extended text online.

As a 'call to action' a QR tag may work better than a printed URL because it saves typing in a URL on a mobile keyboard.

QR tags would also work well as physical world hyperlinks, providing a visible sign that information about a particular location is available online or as a short piece of text encoded in the QR tag. They could work as well for a guerrilla campaign to make contested or forgotten histories visible again - stickers are easy to produce and can be replaced if they weather - as for official projects to take cultural heritage content outside the walls of the museum.

The Powerhouse Museum have also experimented with QR tags, creating special offer vouchers.

Here's the obligatory sample QR - if your phone has a barcode reader you should get the URL of this blog*:

qrcode

* which is totally not optimised for mobile reading as the main pages tend to be quite long but it works ok over wifi broadband.

[Update - I just came across this post about Barcode wikipedia that suggests: "People would be able to access the info by entering/scanning the barcode number. The kind of information that would be stored against the product would be things like reviews, manufacturing conditions, news stories about the product/manufacturer, farm subsidies paid to the manufacturer etc." I'm a bit (ok, a lot) of a hippie and check product labels before I buy - I love this idea because it's like a version of the ethical shopping guide small enough to fit inside my wap phone.]

[Update 2 - more discussion of a 'what are QR codes good for' ilk over at http://blog.paulwalk.net/2008/10/24/quite-resourceful/]

Wednesday, 22 October 2008

Nice summary of web 2.0 for the digital humanities

It's an old post (2006, gasp!) but the points Web 2.0 and the Digital Humanities raises are still just as relevant in the digital cultural heritage sector today:

In summary:
  • Give users tools to visualise and network their own data. And make it easy.
  • Harness the self-interest of your users - "help the user with their own research interests as a first priority".
  • Have an API -"You don’t know what you’ve got until you give it away", "Sharing data in a machine readable and retrievable format, is the most important feature. It lets other people build features for you"
  • Embrace the chaos of knowledge - "a bottom-up method of knowledge representation can be more powerful and more accurate than traditional top-down methods".

Monday, 14 July 2008

Microupdates and you (a.ka. 'twits in the museum')

I was trying to describe Twitter-esque applications for a presentation today, and I wasn't really happy with 'microblogging' so I described them as 'micro-updates'. Partly because I think of them as a bit like Facebook status updates for geeks, and partly because they're a lot more actively social than blog posts.

In case you haven't come across them, Twitter, Pownce, Jaiku, tumblr, etc, are services that let you broadcast short (140 characters) messages via a website or mobile device. I find them useful for finding like-minded people (or just those who also fancy a drink) at specific events (thanks to Brian Kelly for convincing me to try it).

You can promote a 'hash tag' for use at your event - yes, it's a tag with a # in front of it, low tech is cool. Ideally your tag should be short and snappy yet distinct, because it has to be typed in manually (mistakes happen easily, especially from a mobile device) and it's using up precious characters. You can use tools like Summize, hashtags, Quotably or Twemes to see if anyone else has used the same tag recently.

You can also ask people to use your event tag on blog posts, photos and videos to help bring together all the content about your event and create an ad hoc community of participants. Be aware that especially with Twitter-type services you may get fairly direct criticism as well as praise - incredibly useful, but it can seem harsh out of context (e.g. in a report to your boss).

More generally, you can use the same services above to search twitter conversations to find posts about your institution, events, venues or exhibitions. You can add in a search term and subscribe to an RSS feed to be notified when that term is used. For example, I tried http://summize.com/search?q="museum+of+london" and discovered a great review of the last 'Lates' event that described it as 'like a mini festival'. You should also search for common variations or misspellings, though they may return more false positives. When someone tweets (posts) using your search phrase it'll show up in your RSS reader and you can then reply to the poster or use the feedback to improve your projects.

This can be a powerful way to interact with your audience because you can respond directly and immediately to questions, complaints or praise. Of course you should also set up google alerts for blog posts and other websites but micro-update services allow for an incredible immediacy and directness of response.

As an example, yesterday I tweeted (or twitted, if you prefer):
me: does anyone know how to stop firefox 3 resizing pages? it makes images look crappy
I did some searching [1] and found a solution, and posted again:
me: aha, it's browser.zoom.full or "View → Zoom → Zoom Text Only" on windows, my firefox is sorted now
Then, to my surprise, I got a message from someone involved with Firefox [2]:
firefox_answers: Command/Control+0 (zero, not oh) will restore the default size for a page that's been zoomed. Also View->Zoom->Reset

me: Impressed with @firefox_answers providing the answer I needed. I'd been looking in the options/preferences tabs for ages

firefox_answers: Also, for quick zooming in & out use control plus or control minus. in Firefox 3, the zoom sticks per site until you change it.
Not only have I learnt some useful tips through that exchange, I feel much more confident about using Firefox 3 now that I know authoritative help is so close to hand, and in a weird way I have established a relationshp with them.

Finally, twitter et al have a social function - tonight I met someone who was at the same event I was last week who vaguely recognised me because of the profile pictures attached to Twitter profiles on tweets about the event. Incidentally, he's written a good explanation of twitter, so I needn't have written this!

[1] Folksonomies to the rescue! I'd been searching for variations on 'firefox shrink text', 'firefox fit screen', 'firefox screen resize' but since the article that eventually solved my problem called it 'zoom', it took me ages to find it. If the page was tagged with other terms that people might use to describe 'my page jumps, everything resizes and looks a bit crappy' in their own words, I'd have found the solution sooner.

[2] Anyone can create a username and post away, though I assume Downing Street is the real thing.

Saturday, 5 July 2008

Introducing modern bluestocking

[Update, May 2012: I've tweaked this entry so it makes a little more sense.  These other posts from around the same time help put it in context: Some ideas for location-linked cultural heritage projectsExposing the layers of history in cityscapes, and a more recent approach  '...and they all turn on their computers and say 'yay!'' (aka, 'mapping for humanists'). I'm also including below some content rescued from the ning site, written by Joanna:
What do historian Catharine Macauley, scientist Ada Lovelace, and photographer Julia Margaret Cameron have in common? All excelled in fields where women’s contributions were thought to be irrelevant. And they did so in ways that pushed the boundaries of those disciplines and created space for other women to succeed. And, sadly, much of their intellectual contribution and artistic intervention has been forgotten.

Inspired by the achievements and exploits of the original bluestockings, Modern Bluestockings aims to celebrate and record the accomplishments not just of women like Macauley, Lovelace and Cameron, but also of women today whose actions within their intellectual or professional fields are inspiring other women. We want to build up an interactive online resource that records these women’s stories. We want to create a feminist space where we can share, discuss, commemorate, and learn.

So if there is a woman whose writing has inspired your own, whose art has challenged the way you think about the world, or whose intellectual contribution you feel has gone unacknowledged for too long, do join us at http://modernbluestocking.ning.com/, and make sure that her story is recorded. You'll find lots of suggestions and ideas there for sharing content, and plenty of willing participants ready to join the discussion about your favourite bluestocking.
And more explanation from modernbluestocking on freebase:
Celebrating the lives of intellectual women from history...

Wikipedia lists bluestocking as 'an obsolete and disparaging term for an educated, intellectual woman'.  We'd prefer to celebrate intellectual women, often feminist in intent or action, who have pushed the boundaries in their discipline or field in a way that has created space for other women to succeed within those fields.

The original impetus was a discussion at the National Portrait Gallery in London held during the exhibition 'Brilliant Women, 18th Century Bluestockings' (http://www.npg.org.uk/live/wobrilliantwomen1.asp) where it was embarrassingly obvious that people couldn't name young(ish) intellectual women they admired.  We need to find and celebrate the modern bluestockings.  Recording and celebrating the lives of women who've gone before us is another way of doing this.
However, at least one of the morals of this story is 'don't get excited about a project, then change jobs and start a part-time Masters degree.  On the other hand, my PhD proposal was shaped by the ideas expressed here, particularly the idea of mapping as a tool for public history by e.g using geo-located stories to place links to content in the physical location.

While my PhD has drifted away from early scientific women, I still read around the subject and occasionally adding names to modernbluestocking.freebase.com.  If someone's not listed in Wikipedia it's a lot harder to add them, but I've realised that if you want to make a difference to the representation of intellectual women, you need to put content where people look for information - i.e. Wikipedia.

And with the launch of Google's Knowledge Graph, getting history articles into Wikipedia then into Freebase is even more important for the visibility of women's history: "The Knowledge Graph is built using facts and schema from Freebase so everyone who has contributed to Freebase had a part in making this possible. ...The Knowledge Graph is built using facts and schema from Freebase soeveryone who has contributed to Freebase had a part in making this possible. (Source: this post to the Freebase list).  I'd go so far as to say that if it's worth writing a scholarly article on an intellectual woman, it's worth re-using  your references to create or improve their Wikipedia entry.]

Anyway. On with the original post...]

I keep meaning to find the time to write a proper post explaining one of the projects I'm working on, but in the absence of time a copy and paste job and a link will have to do...

I've started a project called 'modern bluestocking' that's about celebrating and commemorating intellectual women activists from the past and present while reclaiming and redefining the term 'bluestocking'.  It was inspired by the National Portrait Gallery's exhibition, 'Brilliant Women: 18th-Century Bluestockings'.  (See also the review, Not just a pretty face).

It will be a website of some sort, with a community of contributors and it'll also incorporate links to other resources.

We've started talking about what it might contain and how it might work at modernbluestocking.ning.com (ning died, so it's at modernbluestocking.freebase.com...)

Museum application (something to make for mashed museum day?): collect feminist histories, stories, artefacts, images, locations, etc; support the creation of new or synthesised content with content embedded and referenced from a variety of sources. Grab something, tag it, display them, share them; comment, integrate, annotate others. Create a collection to inspire, record, commemorate, and build on.
What, who, how should this website look? Join and help us figure it out.

Why modernbluestocking? Because knowing where you've come from helps you know where you're going.

Sources could include online exhibition materials from the NPG (tricky interface to pull records from).  How can this be a geek/socially friendly project and still get stuff done?  Run a Modernbluestocking, community and museum hack day app to get stuff built and data collated?  Have list of names, portraits, objects for query. Build a collection of links to existing content on other sites? Role models and heroes from current life or history. Where is relatedness stored? 'Significance' -thorny issue? Personal stories cf other more mainstream content?  Is it like a museum made up of loan objects with new interpretation? How much is attribution of the person who added the link required? Login v not? Vandalism? How do deal with changing location or format of resources? Local copies or links? Eg images. Local don't impact bandwidth, but don't count as visits on originating site. Remote resources might disappear - moved, permissions changed, format change, taken offline, etc, or be replaced with different content. Examine the sources, look at their format, how they could be linked to, how stable they appear to be, whether it's possible to contact the publisher...

Could also be interesting to make explicit, transparent, the processes of validation and canonisation.

Friday, 6 June 2008

Yahoo! SearchMonkey, the semantic web - an example from last.fm

I had meant to blog about SearchMonkey ages ago, but last.fm's post 'Searching with my co-monkey' about a live example they've created on the SearchMonkey platform has given me the kick I needed. They say:
The first version of our application deals with artist, album and track pages giving you a useful extract of the biography, links to listen to the artist if we have them available, tags, similar artists and the best picture we can muster for the page in question.
Some background on SearchMonkey from ReadWriteWeb:
At the same time, it was clear that enhancing search results and cross linking them to other pieces of information on the web is compelling and potentially disruptive. Yahoo! realized that in order to make this work, they need to incentivize and enable publishers to control search result presentation.

...

SearchMonkey is a system that motivates publishers to use semantic annotations, and is based on existing semantic standards and industry standard vocabularies. It provides tools for developers to create compelling applications that enhance search results. The main focus of these applications is on the end user experience - enhanced results contain what Yahoo! calls an "infobar" - a set of overlays to present additional information.

...

SearchMonkey's aim is to make information presentation more intelligent when it comes to search results by enabling the people who know each result best - the publishers - to define what should be presented and how.
(From Making the Web Searchable: The Story of SearchMonkey)

And from Yahoo!'s search blog:
This new developer platform, which we're calling SearchMonkey, uses data web standards and structured data to enhance the functionality, appearance and usefulness of search results. Specifically, with SearchMonkey:

  • Site owners can build enhanced search results that will provide searchers with a more useful experience by including links, images and name-value pairs in the search results for their pages (likely resulting in an increase in traffic quantity and quality)
  • Developers can build SearchMonkey apps that enhance search results, access Yahoo! Search's user base and help shape the next generation of search
  • Users can customize their search experience with apps built by or for their favorite sites
This could be an interesting new development - the question is, how well does the data we currently output play with it; could we easily adapt our pages so they're compatible with SearchMonkey; should we invest the time it might take? Would a simple increase in the visibility and usefulness of search results be enough? Could there be a greater benefit in working towards federated searches across the cultural heritage sector or would this require a coordinated effort and agreement on data standards and structure?

Update to link to the Yahoo! Search Blog post ;The Yahoo! Search Gallery is Open for Business' which has a few more examples.

Saturday, 31 May 2008

Some ideas for location-linked cultural heritage projects

I loved the Fire Eagle presentation I saw at the WSG Findability event [my write-up] because it got me all excited again about ideas for projects that take cultural heritage outside the walls of the museum, and more importantly, it made some of those projects seem feasible.

There's also been a lot of talk about APIs into museum data recently and hopefully the time has come for this idea. It'd be ace if it was possible to bring museum data into the everyday experience of people who would be interested in the things we know about but would never think to have 'a museum experience'.

For example, you could be on your way to the pub in Stoke Newington, and your phone could let you know that you were passing one of Daniel Defoe's hang outs, or the school where Mary Wollstonecraft taught, or that you were passing a 'Neolithic working area for axe-making' and that you could see examples of the Neolithic axes in the Museum of London or Defoe's headstone in Hackney Museum.

That's a personal example, and those are some of my interests - Defoe wrote one of my favourite books (A Journal of the Plague Year), and I've been thinking about a project about 'modern bluestockings' that will collate information about early feminists like Wollstonecroft (contact me for more information) - but ideally you could tailor the information you receive to your interests, whether it's football, music, fashion, history, literature or soap stars in Melbourne, Mumbai or Malmo. If I can get some content sources with good geo-data I might play with this at the museum hack day.

I'm still thinking about functionality, but a notification might look something like "did you know that [person/event blah] [lived/did blah/happened] around here? Find out more now/later [email me a link]; add this to your map for sharing/viewing later".

I've always been fascinated with the idea of making the invisible and intangible layers of history linked to any one location visible again. Millions of lives, ordinary or notable, have been lived in London (and in your city); imagine waiting at your local bus stop and having access to the countless stories and events that happened around you over the centuries. Wikinear is a great example, but it's currently limited to content on Wikipedia, and this content has to pass a 'notability' test that doesn't reflect local concepts of notability or 'interestingness'. Wikipedia isn't interested in the finds associated with an archaeological dig that happened at the end of your road in the 1970s, but with a bit of tinkering (or a nudge to me to find the time to make a better programmatic interface) you could get that information from the LAARC catalogue.

The nice thing about local data is that there are lots of people making content; the not nice thing about local data is that it's scattered all over the web, in all kinds of formats with all kinds of 'trustability', from museums/libraries/archives, to local councils to local enthusiasts and the occasional raving lunatic. If an application developer or content editor can't find information from trusted sources that fits the format required for their application, they'll use whatever they can find on other encyclopaedic repositories, hack federated searches, or they'll screen-scrape our data and generate their own set of entities (authority records) and object records. But what happens if a museum updates and republishes an incorrect record - will that change be reflected in various ad hoc data solutions? Surely it's better to acknowledge and play with this new information environment - better for our data and better for our audiences.

Preparing the data and/or the interface is not necessarily a project that should be specific to any one museum - it's the kind of project that would work well if it drew on resources from across the cultural heritage sector (assuming we all made our geo-located object data and authority records available and easily queryable; whether with a commonly agreed core schema or our own schemas that others could map between).

Location-linked data isn't only about official cultural heritage data; it could be used to display, preserve and commemorate histories that aren't 'notable' or 'historic' enough for recording officially, whether that's grime pirate radio stations in East London high-rise roofs or the sites of Turkish social clubs that are now new apartment buildings. Museums might not generate that data, but we could look at how it fits with user-generated content and with our collecting policies.

Or getting away from traditional cultural heritage, I'd love to know when I'm passing over the site of one of London's lost rivers, or a location that's mentioned in a film, novel or song.

[Updated December 2008 to add - as QR tags get more mainstream, they could provide a versatile and cheap way to provide links to online content, or 250 characters of information. That's more information than the average Blue Plaque.]

Monday, 26 May 2008

Notes from 'How Can Culture Really Connect? Semantic Front Line Report' at MW2008

These are my notes from the workshop on "'How Can Culture Really Connect? Semantic Front Line Report" at Museums and the Web 2008. This session was expertly led by Ross Parry.

The paper, "Semantic Dissonance: Do We Need (And Do We Understand) The Semantic Web?" (written by Ross Parry, Jon Pratty and Nick Poole) and the slides are online. The blog from the original Semantic Web Think Tank (SWTT) sessions is also public.

These notes are pretty rough so apologies for any mistakes; I hope they're a bit useful to people, even though it's so late after the event. I've tried to include most of what was discussed but it's taken me a while to catch up.

There's so much to see at MW I missed the start of this session; when we arrived Ross had the participants debating the meaning of terms like 'Web 2.0', 'Web 3.0', 'semantic web, 'Semantic Web'.

So what is the semantic web (sw) about? It's about intelligent and efficient searching; discovering resources (e.g. URIs of picture, news story, video, biographical detail, museum object) rather than pages; machine-to-machine linking and processing of data.

Discussion: how much/what level of discourse do we need to take to curators and other staff in museums?
me: we need to show people what it can do, not bother them with acronyms.
Libby Neville: believes in involving content/museum people, not sure viewing through the prism of technology.
[?]: decisions about where data lives have an effect.

Slide 39 shows various axes against which the Semantic Web (as formally defined) and the semantic web (the SW 'lite'?) can be assessed.
Discussion: Aaron: it's context-dependent.

'expectations increase in proportion to the work that can be done' so the work never decreases.

sw as 'webby way to link data'; 'machine processable web' saves getting hung up on semantics [slide 40 quoting Emma Tonkin in BECTA research report, ‘If it quacks like a duck…’ Developments in search technologies].

What should/must/could we (however defined) do/agree/build/try next (when)?

Discussion: Aaron: tagging, clusters. Machine tags (namespace: predicate: value).
me: let's build semantic webby things into what we're doing now to help facilitate the conversations and agreements, provide real world examples - attack the problem from the bottom up and the top down.

Slide 49 shows three possible modes: make collections machine-processable via the web; build ontologies and frameworks around added tags; develop more layered and localised meaning. [The data (the data around the data) gets smarter and richer as you move through those modes.]

I was reminded of this 'mash it' video during this session, because it does a good jargon-free job of explaining the benefits of semantic webby stuff. I also rather cynically tweeted that the semantic web will "probably happen out there while we talk about it".

Sunday, 20 April 2008

Crowdsourcing metadata cleaning?

If you're interested in another perspective on dealing with user-generated tags or metadata, this blog post from last.fm, Fingerprinting and Metadata Progress Report talks about how they're trying to create 'order from chaos':
So far our fingerprint server identified 23 million unique tracks, from the 650 million fingerprint requests you’ve thrown at it. Who knows how many unique tracks there are out there.. We have a couple of hundred million tracks based on spelling alone – but not all of them are spelt correctly.

They have some interesting issues to deal with in cleaning up their (i.e. your data, if you're a last.fm user) data, especially when 'the most popular spelling is not necessarily the correct one'. And what about bands that change their name (but are essentially the same band) or line-up (are they still the same band?) - when do you decide to create a new identifier?

They're letting users who are logged in vote on potential corrections to an artist name, effectively testing crowdsourcing metadata corrections as well as the original data creation process. This model could work for museums - depending on the collection, some museums already get a lot of corrections when parts of their collections are published online. What would happen if we made that process transparent?

Saturday, 19 April 2008

Museums and Clayton's audience participation

A comment Seb left on Nate's blog post about "master" metadata got me thinking about cognitive dissonance and whether museums who say they're open to public participation and content really act as if they are. Are we providing a Clayton's call for audience participation?

If what you do - raise the barrier to participation so high that hardly anyone is going to bother commenting or tagging - speaks louder than what you say - 'sure, we'd love to hear what you have to say' - which one do you think wins?

To pick an example I've seen recently (and this is not meant to be a criticism of them or their team because I have no idea what the reasons were) the London Transport Museum have put 'all Museum objects and stories on display in the new Museum' on their collections website, which is fantastic. If you look at a collection item, the page says, "Share a story with us - comment on this image", which sounds really open and inviting.

But
, if you want to comment, they ask for a lot of information about you first - check this random example.

So, ok. There are lots of possible reasons for this. UK museums have to deal with the Data Protection Act, which might complicate things, and their interpretation of the DPA might mean they ask for more information rather than less and add that scary tick box.

Or maybe they think the requirement to give this information won't deter their audience. I'd imagine that London Transport Museum's specialist audiences won't be put off by a registration form - some of their users are literally trainspotters and at risk of believing a stereotype, if they can bear the kind of weather that requires anoraks, they're probably not put off by a form.

Or maybe they're trying to control spam (though email addresses are no barrier to spam, and it's easy to use Akismet or moderation to trap spam); or maybe it's a halfway house between letting go and keeping control; or maybe they're tweaking the form in response to usage and will lower the barriers if they're not getting many comments.

Or maybe it's because the user-generated content captured this way goes directly into their collection management system and they want to record some idea of the provenance of the data. From a post to the UK Museums Computer Group list:
We have just launched the London Transport Online Museum. Users can view
every object, gallery and label text on display in our new museum in Covent Garden.

Following on from the current discussion thread we have incorporated into this new site, the facility for users to leave us memories / stories on all objects on display. Rather than a Wiki submission these stories are made directly on the website and will be fed back into our collection management system. These submissions can be viewed by all users as soon as they have passed through moderation process.

We will closely monitor how many responses we get and feedback to the group.

Please have a look, and maybe even leave us a memory?
[My emphasis in bold]

Moving on from the example of the London Transport Museum...

Whether the gap between their stated intentions and the apparent barriers to accepting user-generated content is the result of internal ambivalence about or resistance to user-generated content, concern about spam or 'bad data', or a belief that their specialist audiences will persist despite the barriers doesn't really make a difference; ultimately the intentionality matters less than the effect.

By raising the barrier to participation, aren't they ensuring that the casual audience remains exactly that - interested, but not fully engaged?

And as Seb pointed out, "Remembering that even tagging on the PHM collection - 15million views in 2007, 5 thousand tags . . . - and that is without requiring ANY form of login."

It also reminds me of what Peter Samis said at Museums and the Web in Montreal about engaging with museum visitors digitally: "We opened the door to let visitors in... then we left the room".

(If you're curious, the title is a reference to an Australian saying: Clayton's was "the drink you have when you're not having a drink", as as Wikipedia has it 'a compromise which satisfies no-one'. 'Ersatz' might be another word for it.)

Wednesday, 16 April 2008

Questions from 'Beyond Single Repositories' at MW2008

I'm still working on getting my notes from Museums and the Web in Montreal online.

These are notes from the questions at the 'Beyond Single Repositories' session. This session was led by Ross Parry, and included the papers Learning from the People: Traditional Knowledge and Educational Standards by Daniel Elias and James Forrest and The Commons on Flickr: A Primer by George Oates.

This clashed with the User-Generated Content session that I felt I should see for work, but I managed to sneak in at the end of Ross's session. I expected this room to be packed, but it wasn't. I guess the ripples of user-generated content and Web 2.0-ish stuff are still spreading beyond the geeks, and the pebbles of single repositories and the semantic web have barely dropped into the pond for most people. As usual, all mistakes are mine - if you asked a question and I haven't named you or got your question wrong, drop me a line.

Quite a lot of the questions related to 'The Commons'.

There was a question about the difference between users who download and retain context of images, versus those who just download the image and lose all context, attribution, etc. George: Flickr considered putting the metadata into EXIF but it was problematic and wasn't robust enough to be useful.

Another question: how to link back to institution from Flickr? George: 'there's this great invention called the hyperlink'. And links can also go to picture libraries to buy prints.

[I need to check this but it could really help make the case for Commons in museums if that's the case. We might also be able to target different audiences with different requirements - e.g. commercial publications vs school assignments. I also need to check if Flickr URLs are permanent and stable.]

Seb Chan asked: how does business model of having images on Flickr co-exist with existing practices?

Flickr are cool with museums putting in content at different resolutions - it's up to institution to decide.

"It's so easy to do things the correct way" so please teach everyone to use CC licence stuff appropriately.

Issues are starting to be raised about revenue sharing models.

[I wonder if we could put in FOI requests to find out exactly how much revenue UK museums make from selling images compared to the overhead in servicing commercial picture libraries, and whether it varies by type of image or use. It'd be great if we could put some Museum of London/MoLAS images on Commons, particularly if we could use tagging to generate multilingual labels and re-assess images in terms of diversity - such an important issue for our London audiences; or to get more images/objects geo-located. I also wonder if there are any resourcing issues for moderation requirements, or do we just cope with whatever tags are added?]

Update: following the conference, Frankie Roberto started a discussion on the Museums Computer Group list under the subject 'copyright licensing and museums'. You have to be a member to post but a range of perspectives and expertise would really help move this discussion on.

Thursday, 25 October 2007

Feeds for beginners

From A Consuming Experience, Feeds basics 101: introduction to newsfeeds:

Feeds, RSS feeds, Atom feeds, XML feeds, newsfeeds, web feeds, they're increasingly common on the internet these days - but what are they, how do you subscribe, and how do you publish and publicise your own news feed? This post is a 3-part introductory tutorial guide to web feeds, aimed at intelligent non-geeks

Friday, 21 September 2007

Some random links...

Two very handy resources when choosing forum software: opensourcecms.com lets you try out various installations - you can create test forums and play with the settings and forummatrix.org helps you compares applications on a variety of facets, and there's a wizard to help you narrow the choices.

Andy Powell makes the excellent point that social software-style tags function as virtual venues:
if you are holding an event, or thinking about holding an event... decide what tag you are going to use as soon as possible. ... In fact, in a sense, the tag becomes the virtual venue for the event's digital legacy.


In other news, Intel get into Mashups for the Masses - "an extension to your existing web browser that allows you to easily augment the page that you are currently browsing with information from other websites. As you browse the web, the Mash Maker toolbar suggests Mashups that it can apply to the current page in order to make it more useful for you" and the BBC reports on Metaplace, a "free tool that allows anyone to create a [3D] virtual world" and incorporates lots of social web tools.

Wednesday, 19 September 2007

According to Merlin on the Web: the British Museum Collection Database Goes Public on the CHArt conference site, the British Museum are putting their entire catalogue online. The evaluation will make fascinating reading and I hope can be made public - I'd like to know who uses it and why, does the variation in detail and 'quality' of records matter to them, how much of the collection is accessed, how many corrections or requests for more information are made, and how public comments work in practice.

Sunday, 24 June 2007

Who's talking about you?

This article explains how you can use RSS feeds to track mentions of your company (or museum) in various blog search sites: Ego Searches and RSS.

It's a good place to start if you're not sure what people are saying about your institution, exhibitions or venues or whether they might already be creating content about you. Don't forget to search Flickr and YouTube too.

Thursday, 5 April 2007

I'm sure this has been everywhere already but the NY Times have an excellent article on museums and tagging.

Monday, 12 March 2007

Exposing the layers of history in cityscapes

I really liked this talk on "Time, History and the Internet" because it touches on lots of things I'm interested in.

I have a on-going fascination with the idea of exposing the layers of history present in any cityscape.

I'd like to see content linked to and through particular places, creating a sense of four dimensional space/time anchored specifically in a given location. Discovering and displaying historical content marked-up with the right context (see below) gives us a chance to 'move' through the fourth dimension while we move through the other three; the content of each layer of time changing as the landscape changes (and as information is available).

Context for content: when was it written? Was it written/created at the time we're viewing, or afterwards, or possibly even before it about the future time? Who wrote/created it, and who were they writing/drawing/creating it for? If this context is machine-readable and content is linked to a geo-reference, can we generate a representation of these layers on-the-fly?

Imagine standing at the base of Centrepoint at London's Tottenham Court Road and being able to ask, what would I have seen here ten years ago? fifty? two hundred? two thousand? Or imagine sitting at home, navigating through layers of historic mapping and tilting down from a birds eye view to a view of a street-level reconstructed scene. It's a long way off, but as more resources are born or made discoverable and interoperable, it becomes more possible.

Friday, 2 February 2007

Tagging goes mainstream?

"A December 2006 survey has found that 28% of internet users have tagged or categorized content online such as photos, news stories or blog posts. On a typical day online, 7% of internet users say they tag or categorize online content.
...
Tagging is gaining prominence as an activity some classify as a Web 2.0 hallmark in part because it advances and personalizes online searching. Traditionally, search on the web (or within websites) is done by using keywords. Tagging is a kind of next-stage search phenomenon – a way to mark, store, and then retrieve the web content that users already found valuable and of which they want to keep track. It is, of course, more tailored to individual needs and not designed to be the all-inclusive system"
Pew Internet and American Life project: Tagging

The report also goes into the definition of tagging as well as who tags and there's an interview with David Weinberger on 'Why Tagging Matters'.