Saturday, 1 October 2011
Usability: the key that unlocks geeky goodness
I posted on the project blog about how I worked out a testing plan to encourage user-centred design and set up the usability sessions in Evaluating Pelagios' usability, set out how a test session runs (with sample scripts and tasks) in Evaluating usability: what happens in a user testing session? and finally I posted some early Pelagios usability testing results. The results are from a very small sample of potential users but they were consistent in the issues and positive results uncovered.
The wider lesson for LOD-LAM (linked open data in library, archives, museums) projects is that user testing (and/or a strong user-centred design process) helps general audiences (including subject specialists) appreciate the full potential of a technically-led project - without thoughtful design, the results of all those hours of code may go unloved by the people they were written for. In other words, user experience design is the key that unlocks the geeky goodness that drives these projects. It's old news, but the joy of user testing is that it reminds you of what's really important...
Wednesday, 3 November 2010
What would Phar Lap do? AKA, what happens when Facebook and museum URIs meet a dead horse?
I've always been fascinated by the way the public respond to Phar Lap - when I worked at Museum Victoria, the outreach team would regularly get emails written to Phar Lap by people who had seen the film or somehow come across his story. (I was also never quite sure why they thought emailing a dead horse would work). So when I first heard that Phar Lap was on Facebook, I was curious to see which museum would have 'claimed' Phar Lap. Does possession of the most charismatic object (the hide) make it easier for Melbourne Museum to step up as the presence of Phar Lap on social media, or were they just the first to be in that space? The issues around 'ownership' and right to speak for an iconic object like Phar Lap make a brilliant case study for how museums represent their collections online.
And today, when I came across three posts (Responses to "Progress on Museum URIs", Progress on Museum URIs by @sebastianheath, Identifing Objects in Museum Collections by @ekansa) on movements towards stable museum URIs that problematised the "politics of naming and identifying cultural heritage" and the concept of the "exclusive right of museums to identify their objects", I thought of Phar Lap. (Which is nice, cos 80 years and one day ago he won the Melbourne Cup).
Of the three museums that own bits of the dead horse, which gets to publish the canonical digital record about Phar Lap? I hope the question sounds silly enough to highlight the challenges and opportunities in translating physical models to the digital realm. Of course each museum can publish a record (specifically, mint a URI) about Phar Lap (and I hope they do) but none of the museums could prevent the others from publishing (and hopefully they wouldn't want to).
Or as the various blog posts said, "many agents can assert an identity for an object, with those identities together forming a distributed and diverse commentary on the human past", and museums need to play their part: "a common identifier promoted by and discoverable at the holding institution will ease the process of recognizing that two or more identifiers refer to the 'same thing'".
Of course it's not that simple, and if you're interested in the questions the museum sector (by which I hopefully don't only mean me) is grappling with, the museums and the machine-processable web page on Permanent IDs has links to discussions on the MCG list, and I've wrestled a bit with how URIs might look at the Science Museum/NMSI (and I need to go back and review the comments left by various generous people). I'd love to know what other museums are planning, and what consumers of the data might need, so that we can come up with a robust common model for museum URIs.
And to reward you for getting this far, here is a picture of Phar Lap on Facebook as his skin and bones are about to be re-united:
Sunday, 4 July 2010
Linking museums: machine-readable data in cultural heritage - meetup in London July 7
As I posted to the MCG list: "A very informal meetup to discuss 'Linking museums: machine-readable data in cultural heritage' is happening next Wednesday. I'm hoping for a good mix of people with different levels of experience and different perspectives on the issue of publishing data that can be re-used outside the institution that created it. ... please do pass this on to others who may be interested. If you would like to come but can't get down to that London, please feel free to send me your questions and comments (or beer money)."
The basic details are: July 7, 2010, Shooting Star pub, London. 7:30 - 10pm-ish. More information is available at http://museum-api.pbworks.com/July-2010-meetup and you can let me know you're coming or register your interest.
In more detail...
Why?
I'm trying to cut through the chicken and egg problem - as a museum technologist, I can work towards getting machine-readable data available, but I'm not sure which formats and what data would be most useful for developers who might use it. Without a critical mass of take-up for any one type, the benefits of any one data source are more limited for developers. But museums seem to want a sense of where the critical mass is going to be so they can build for that. How do we cut through this and come up with a sensible roadmap?
Who?
You! If you're interested in using museum data in mashups but find it difficult to get started or find the data available isn't easily usable; if you have data you want to publish; if you work in a museum and have a
data publication problem you'd like help in solving; if you are a cheerleader for your favourite acronym...
Put another way, this event is for you if you're interested in publishing and sharing data about their museums and collections through technologies such as linked data and microformats.
It'll be pretty informal! I'm not sure how much we can get done but it'd be nice to put faces to names, and maybe start some discussions around the various problems that could be solved and tools that could be
created with machine-readable data in cultural heritage.
Sunday, 21 March 2010
Some thoughts on linked data and the Science Museum - comments?
Wednesday, 13 May 2009
Final thoughts on open hack day (and an imaginary curatr)
I went because it tied in really well with some work projects (like the museum metadata mashup competition we're running later in the year or the attempt to get a critical mass of vaguely compatible museum data available for re-use) and stuff I'm interested in personally (like modern bluestocking, my project for this summer - let me know if you want to help, or just add inspiring women to freebase).
I'm also interested in creating something like a Dopplr for museums - you tell it what you're interested in, and when you go on a trip it makes you a map and list of stuff you could see while you're in that city.
Like: I like Picasso, Islamic miniatures, city museums, free wine at contemporary art gallery openings, [etc]; am inspired by early feminist history; love hearing about lived moments in local history of the area I'll be staying in; I'm going to Barcelona.
The 'list of cultural heritage stuff I like' could be drawn from stuff you've bookmarked, exhibitions you've attended (or reviewed) or stuff favourited in a meta-museum site.
(I don't know what you'd call this - it's like a personal butlr or concierge who knows both your interests and your destinations - curatr?)
The talks on RDFa (and the earlier talk on YQL at the National Maritime Museum) have inspired me to pick a 'good enough' protocol, implement it, and see if I can bring in links to similar objects in other museum collections. I need to think about the best way to document any mapping I do between taxonomies, ontologies, vocabularies (all the museumy 'ies') and different API functions or schemas, but I figure the museum API wiki is a good place to draft that. It's not going to happen instantly, but it's a good goal for 2009.
These are the last of my notes from the weekend's Open Hack London event, my notes from various talks are tagged openhacklondon.
Tom Morris, SPARQL and semweb stuff - tech talk at Open Hack London
He's since posted his links and queries - excellent links to endpoints you can test queries in.
Semantic web often thought of as long-promised magical elixir, he's here to say it can be used now by showing examples of queries that can be run against semantic web services. He'll demonstrate two different online datasets and one database that can be installed on your own machine.
First - dbpedia - scraped lots of wikipedia, put it into a database. dbpedia isn't like your averge database, you can't draw a UML diagram of wikipedia. It's done in RDF and Linked Data. Can be queried in a language that looks like SQL but isn't. SPARQL - is a w3c standard, they're currently working on SPARQL 2.
Go to dbpedia.org/sparql - submit query as post. [Really nice - I have a thing about APIs and platforms needing a really easy way to get you to 'hello world' and this does it pretty well.]
[Line by line comments on the syntax of the queries might be useful, though they're pretty readable as it is.]
'select thingy, wotsit where [the slightly more complicated stuff]'
Can get back results in xml, also HTML, 'spreadsheet', JSON. Ugly but readable. Typed.
[Trying a query challenge set by others could be fun way to get started learning it.]
One problem - fictional places are in Wikipedia e.g. Liberty City in Grand Theft Auto.
Libris - how library websites should be
[I never used to appreciate how much most library websites suck until I started back at uni and had to use one for more than one query every few years]
Has a query interface through SPARQL
Comment from the audience BBC - now have SPARQL endpoint [as of the day before? Go BBC guy!].
Playing with mulgara, open source java triple store. [mulgara looks like a kinda faceted search/browse thing] Has own query language called TQL which can do more intresting things than SPARQL. Why use it? Schemaless data storage. Is to SQL what dynamic typing is to static typing. [did he mean 'is to sparql'?]
Question from audence: how do you discover what you can query against?
Answer: dbpedia website should list the concepts they have in there. Also some documentation of categories you can look at. [Examples and documentation are so damn important for the update of your API/web service.]
Coming soon [?] SPARUL - update language, SPARQL2: new features
The end!
[These are more (very) rough notes from the weekend's Open Hack London event - please let me know of clarifications, questions, links or comments. My other notes from the event are tagged openhacklondon.
Quick plug: if you're a developer interested in using cultural heritage (museums, libraries, archives, galleries, archaeology, history, science, whatever) data - a bunch of cultural heritage geeks would like to know what's useful for you (more background here). You can comment on the #chAPI wiki, or tweet @miaridge (or @mia_out). Or if you work for a company that works with cultural heritage organisations, you can help us work better with you for better results for our users.]
There were other lightning talks on Pachube (pronounced 'patchbay', about trying to build the internet of things, making an API for gadgets because e.g. connecting hardware to the web is hard for small makers) and Homera (an open source 3d game engine).
Saturday, 9 May 2009
Rasmus Lerdorf on Hacking with PHP - tech talk at Open Hack London
Hacking with PHP, Rasmus Lerdorf
Goal of talk: copy and pastable snippets that just work so you don't have to fight to get things that work [there's not enough of this to help beginners get over that initial hump]. The slides are available at http://talks.php.net/show/openhack and these notes are probably best read as commentary alongside the code examples.
[Since it's a hack day, some] Hack ideas: fix something you use every day; build your own targeted search engine; improve the look of search results; play with semantic web tools to make the web more semantic; tell the world what kind of data you have - if a resume, use hResume or other appropriate microformats/markup; go local - tools for helping your local community; hack for good - make the world a better place.
SearchMonkey and BOSS are blending together a little bit.
What we need to learn
With PHP – enough to handle simple requests; talk to backend datastore; how to parse XML with PHP, how to generate JSON, some basic javasccript, a JavaScript utility library like YUI or jquery.
parsing XML: simpleXML_load_file() - can load entire URL or local file.
Attributes on node show up as array. Namespace attributes call children of node, name namespace as argument.
Now know how to parse XML, can get lots of other stuff.
Context extraction service, Yahoo - doesn't get enough attention. Post all text, gives you back four or five key terms - can then do an image search off them. Or match ads to webpages.
Can use get or post (curl) - usually too much for get.
PHP to JavaScript on initial page load: JSON_encode -> javascript.
Javascript to PHP (and back)
If you can figure out these six lines of code, you can write anything in the world. How every modern web application works.
Server-side php, client-side javascript.
'There's nothing to building web applications, you just have to break everything down into small enough chunks that it all becomes trivial'.
AJAX in 30 seconds.
Inline comments in code would help for people reading it without hearing the talk at the same time.
JavaScript libraries to the rescue
load maps API, create container (div) for the map, then fill it.
Form - on submit call return updateMap(); with new location.
YGeoRSS - if have GeoRSS file... can point to it.
GeoPlanet - assigns a WOE ID to a place. Locations are more than just a lat long - carry way more information. Basically gives you a foreign key. YQL is starting to make the web a giant database. Can make joins across APIs - woeid works as fk.
YQL - 'combines all the APIs on the web into a single API'.
Add a cache - nice to YQL, and also good for demos etc. Copy and paste cache function from his slides - does a local cache on URL. Hashed with md5. Using PHP streams - #defn. Adding a cache speeds up developing when hacking (esp as won't be waiting for the wifi). [This is a pretty damn good tip cos it's really useful and not immediately obvious.]
XPath on URL using PHP's OAuth extension
SearchMonkey - social engineering people into caring about semantic data on the web. For non-geeks, search plug-in mechanism that will spruce up search results page. Encourages people to add semantic data so their search result is as sexy as their competitors - so goal is that people will start adding semantic data.
'If you're doing web stuff, and don't know about microformats, and your resume doesn't have hResume, you're not getting a job with Yahoo.'
Question: how are microformats different to RDFa?
Answer: there are different types of microformats - some very specific ones, eg hResume, hCal. RDFa - adding arbitrary tags to page. even if no specific way to describe your data. But there's a standard set of mark-ups for a resume so can use that. if your data doesn't match anything at microfomats.org then use RDFa or erdf (?).
RDFa, SearchMonkey - tech talks at Open Hack London
I'm putting my rough and ready notes online so that those who couldn't make it can still get some of the benefits. Apologies for any mishearings or mistakes in transcription – leave me a comment with any questions or clarifications.
One of the reasons I was going was to push my thinking about the best ways to provide API-like access to museum information and collections, so my notes will reflect that but I try to generalise where I can. And if you have thoughts on what you'd like cultural heritage institutions to do for developers, let us know! (For background, here's a lightning talk I did at another hack event on happy museums + happy developers = happy punters).
RDFa - now everyone can have an API.
Mark Birkbeck
Going to cover some basic mark-up, and talk about why RDFa is a good thing. [The slides would be useful for the syntax examples, I'll update if they go online.]
RDFa is a new syntax from W3C - a way of embedding metadata (RDF) in HTML documents using attributes.
e.g. <span property="dc:title"> - value of property is the text inside the span.
Because it's inline you don't need to point to another document to provide source of metadata and presentation HTML.
One big advance is that can provide metadata for other items e.g. images, so you can e.g. attach licence info to the image rather than page it's in – e.g. <img src="" rel="licence" resource="[creative commons licence]">
Putting RDFa into web pages means you've now got a feed (the web page is the RSS feed), and a simple static web page can become an API that can be consumed in the same way as stuff from a big expensive system. 'Growing adoption'.
Government department Central Office of Information [?] is quite big on RDFa, have a number of projects with it. [I'd come across the UK Civil Service Job Service API while looking for examples for work presentations on APIs.]
RDFa allows for flexible publishing options. If you're already publishing HTML, you can add RDFa mark-up then get flexible publishing models - different departments can keep publishing data in their own way, a central website can go and request from each of them and create its own database of e.g. jobs. Decentralised way of approaching data distribution.
Can be consumed by: smarter browsers; client-side AJAX, other servers such as SearchMonkey.
He's interested where browsers can do something with it - either enhanced browsers that could e.g. store contact info in a page into your address book; or develop JavaScript libraries that can parse page and do something with it. [screen shot of jobs data in search monkey with enhanced search results]
RDFa might be going into Drupal core.
Example of putting isbn in RDFa in page, then a parser can go through the page, pull out the triples [some explanation of them as mini db?], pull back more info about the book from other APIs e.g. Amazon - full title, thumbnail of cover. e.g. pipes.
Example of FOAF - twitter account marked up in page, can pull in tweets. Could presumably pull in newer services as more things were added, without having to re-mark-up all the pages.
Example of chemist writing a blog who mentions a chemical compound in blog post, a processor can go off and retrieve more info - e.g. add icon for mouseover info - image of molecule, or link to more info.
Next plan is to link with BOSS. Can get back RDFa from search results - augment search results with RDFa from the original page.
Search Monkey (what it is and what you can do with it)
Neil Crosby (European frontend architect for search at Yahoo).
SearchMonkey is (one of) Yahoo's open search platforms (along with BOSS). Uses structured data to enhance search results. You get to change stuff on Yahoo search results page.
SearchMonkey lets you: style results for certain URL patterns; brand those results; make the results more useful for users.
[examples of sites that have done it to see how their results look in Yahoo? I thought he mentioned IMDb but it doesn't look any different - a film search that returns a wikipedia result, OTOH, does.]
Make life better for users - not just what Yahoo thinks results should be, you can say 'actually this is the important info on the page'
Three ways to do it [to change the SERP [search engine results page]: mark up data in a way that Yahoo knows about - 'just structure your data nicely'. e.g. video mark-up; enhance a result directly; make an infobar.
Infobar - doesn't change result see immediately on the page, but it opens on the page. e.g. of auto-enhanced result- playcrafter. Link to developer start page - how to mark it up, with examples, and what it all means.
User-enhanced result - Facebook profile pages are marked up with microformats - can add as friend, poke, send message, view friends, etc from the search results page. Can change the title and abstract, add image, favicon, quicklinks, key/value pairs. Create at [link I can't see but is on slides] Displayed in screen, you fill it out on a template.
Infobar - dropdown in grey bar under results. Can do a lot more, as it's hidden in the infobar and doesn't have to worry people.
Data from: microformats, RDF, XSLT, Yahoo's index, and soon, top tags from delicious.
If no machine data, can write an XSLT. 'isn't that hard'. Lots of documentation on the web.
Examples of things that have been made - a tool that exposes all the metadata known for a page. URL on slide. can install on Yahoo search page, add it in. Use location data to make a map - any page on web with metadata about locations on it - map monkey. Get qype results for anything you search for.
There's a mailing list (people willing and wanting to answer questions) and a tutorial.
Questions
Question: do you need to use a special doctype [for RDFa]?
Answer: added to spec that 'you should use this doctype' but the spec allows for RDFa to be used in situations when can't change doctype e.g. RDFa embedded in blogger blogpost. Most parsers walk the DOM rather than relying on the doctype.
Jim O'D - excited that SearchMonkey supports XSLT - if have website with correctly marked up tables, could expose those as key/value pairs?
Answer: yes. XSLT fantastic tool for when don't have data marked up - can still get to it.
Frankie - question I couldn't hear. About info out to users?
Answer: if you've built a monkey, up to you to tell people about it for the moment. Some monkeys are auto-on e.g. Facebook, wikipedia... possibly in future, if developed a monkey for a site you own, might be able to turn it auto-on in the results for all users... not sure yet if they'll do it or not.
Frankie: plan that people get monkeys they want, or go through gallery?
Answer: would be fantastic if could work out what people are using them for and suggest ones appropriate to people doing particular kinds of searches, rather than having to go to a gallery.
Sunday, 29 March 2009
Tim Berners-Lee at TED on 'database hugging' and linked data
I've put some notes below - I was transcribing it for myself and thought I might as well share it. It's only a selection of the talk and I haven't tidied it because they're not my words to edit.
Why is linked data important?
Making the world run better by making this data available. If you know about some data in some government department you often find that, these people, they're very tempted to keep it, to hug your database, you don't want to let it go until you've made a beautiful website for it. ... Who am I to say "don't make a website..." make a beautiful website, but first, give us the unadulterated data. Give us the raw data now.
You have no idea, the number of excuses people come up with to hang onto their data and not give it to you, even though you've paid for it.
Communicating science over the web... the people who are going to solve those are scientists, they have half-formed ideas in their head, but a lot of the state of knowledge of the human race at the moment is in database, currently not sharing. Alzheimer's scientists ... the power of being able ask questions which bridge across different disciplines is really a complete sea-change, it's very, very important. Scientists are totally stymied at the moment, the power of the data that other scientists have collected is locked up, and we need to get it unlocked so we can tackle those huge problems. if I go on like this you'll think [all data from] huge institutions but it's not. [Social networking is data.]
Linked data is about people doing their bit to produce their bit, and it all connecting. That's how linked data works. ... You do your bit, everybody else does theirs. You may not have much data yourself, to put on there, but you know to demand it.
It's not just about the number of places where data comes. It's about connecting it together. When you connect it together you get this power... out of it. It'll only really pay off when everybody else has done it. It's called Linked Data, I want you to make it, I want you to demand it.
Wednesday, 27 August 2008
User-generated mashups in natural language?
It's a framework that brings together lots of the bits of functionality that are available with browser extensions and bookmarklets and lets the user run them with natural language commands. One of the goals is to "enable on-demand, user-generated mashups with existing open Web APIs. (In other words, allowing everyone–not just Web developers–to remix the Web so it fits their needs, no matter what page they are on, or what they are doing.)".
It's a long way from being ubiquitous, but it does show that it's increasingly worth publishing your data in re-usable formats. They show an example of address being picked up from microformats in apartment listings and mapped for the user - that kind of mashup was possible before and they're a huge step forward in themselves, but how many users have the skills and time to do it? Being able to use natural language to pull together and use data could bring mash-ups to the general public in a massive way.
Tuesday, 29 July 2008
One step closer to intelligent searching?
Called Cuil [pronounced 'cool'], from the Gaelic for knowledge and hazel, its founders claim it does a better and more comprehensive job of indexing information online.From the Cuil FAQ:
The technology it uses to index the web can understand the context surrounding each page and the concepts driving search requests, say the founders.
But analysts believe the new search engine, like many others, will struggle to match and defeat Google.
...
Instead of just looking at the number and quality of links to and from a webpage as Google's technology does, Cuil attempts to understand more about the information on a page and the terms people use to search. Results are displayed in a magazine format rather than a list.
So Cuil searches the Web for pages with your keywords and then we analyze the rest of the text on those pages. This tells us that the same word has several different meanings in different contexts. Are you looking for jaguar the cat, the car or the operating system?They also provide 'drill-downs' on the results page.
We sort out all those different contexts so that you don't have to waste time rephrasing your query when you get the wrong result.
Different ideas are separated into tabs; we add images and roll-over definitions for each page and then make suggestions as to how you might refine your search. We use columns so you can see more results on one page.
Cuil will direct you to this additional information. By looking at these suggestions, you may discover search data, concepts, or related areas of interest that you hadn’t expected. This is particularly useful when you are researching a subject you don't know much about and aren't sure how to compose the "right" query to find the information you need.I haven't used it enough to work out exactly how it differentiates concepts (tabs) and 'additional information' (drill-downs/categories).
It does a good job on something like the Cutty Sark. Under 'Explore by Category' it offered:
- Buildings And Structures In Greenwich
- Sailboat Names
- Museums In London
- Neighbourhoods Of Greenwich
- School Ships
It didn't do as well with 'samian ware' - the categories picked up all sorts of places and peoples, (and randomly 'American Films'), but while the search results all say that it's 'a kind of bright red Roman pottery' that's not reflected in the categories. Fair enough, there may not be enough information easily available online so that 'Types of Roman pottery' registers as a category.
Incidentally, most of the results listed for 'samian ware' are just recycled entries from Wikipedia. It's a shame the results aren't filtered to remove entries that have just duplicated Wikipedia text. The FAQ says they don't index duplicate content I guess the overall site or page is just different enough to be retained.
It might take a while for museum content to appear in the most useful ways, but it looks like it might be a useful search engine for niche content. From the FAQ again:
We've found that a lot of Web pages have been designed with a small audience in mind—perhaps they are blogs or academic papers with specific interests or pages with family photos. We think that even though these pages aren't necessarily for a wide audience, they contain content that one day you might need.It's all sounding a bit semantic web-ish (and quite a bit 'reacting to Google-ish') and I'll use it for a while to see how it compared to Google. The webmaster information doesn't give any indication of how you could mark up content so the relationships between terms in different contexts is clear, but I guess nice semantic markup would help.
Our job is to index all these pages and examine their content for relevancy to your search. If they contain information you need, then they should be available to you.
Refreshingly, it doesn't retain search info - privacy is one of their big differentiators from Google.
Tuesday, 8 July 2008
The Future of the Web with Sir Tim Berners-Lee @ Nesta
My notes from the Nesta event, The Future of the Web with Sir Tim Berners-Lee, held in London on July 8, 2008.
As usual, let me know of any errors or corrections, comments are welcome, and comments in [square brackets] are mine. I wanted to get these notes up quickly so they're pretty much 'as is', and they're pretty much about the random points that interested me and aren't necessarily representative. I've written up more detailed notes from a previous talk by Tim Berners-Lee in March 2007, which go into more detail about web science.
[Update: the webcast is online at http://www.nesta.org.uk/future-of-web/ so you might as well go watch that instead.]
The event was introduced by NESTA's CEO, Jonathan Kestenbaum. Explained that online contributions from the pre-event survey, and from the (twitter) backchannel would be fed into the event. Other panel members were Andy Duncan from Channel 4 and the author Charlie Leadbeater though they weren't introduced until later.
Tim Berners-Lee's slides are at http://www.w3.org/2008/Talks/0708-ws-30min-tbl/.
So, onto the talk:
He started designing the web/mesh, and his boss 'didn't say no'.
He didn't want to build a big mega system with big requirements for protocols or standards, hierarchies. The web had to work across boundaries [slide 6?]. URIs are good.
The World Wide Web Consortium as the point where you have to jump on the bob sled and start steering before it gets out of control.
Producing standards for current ideas isn't enough; web science research is looking further out. Slide 12 - Web Science Research Initiative (WSRI) - analysis and synthesis; promote research; new curriculum.
Web as blockage in sink - starts with a bone, stuff builds up around it, hair collect, slime - perfect for bugs, easy for them to get around - we are the bugs (that woke people up!). The web is a rich environment in which to exist.
Semantic web - what's interesting isn't the computers, or the documents on the computers, it's the data in the documents on the computers. Go up layers of abstraction.
Slide on the Linked Open Data movement (dataset cloud) [Anra from Culture24 pointed out there's no museum data in that cloud].
Paraphrase, about the web: 'we built it, we have a duty to study it, to fix it; if it's not going to lead to the kind of society we want, then tweak it, fix it'.
'Someone out there will imagine things we can't imagine; prepare for that innovation, let that innovation happen'. Prepare for a future we can't imagine.
End of talk! Other panelists and questions followed.
Charles Leadbeater - talked about the English Civil War, recommends a book called 'The World Turned Upside Down'. The bottom of society suddenly had the opportunity to be in charge. New 'levellers' movement via the web. Participate, collaborate, (etc) without the trappings of hierarchy. 'Is this just a moment' before the corporate/government Restoration? Iterative, distributed, engaged with practice.
Need new kinds of language - dichotomies like producer/consumer are disabling. Is the web - a mix of academic, geek, rebel, hippie and peasant village cultures - a fundamentally different way of organising, will it last? Are open, collaborative working models that deliver the goals possible? Can we prevent creeping re-regulation that imposes old economics on the new web? e.g. ISPs and filesharing. Media literacy will become increasingly important. His question to TBL - what would you have done differently to prevent spam while keeping the openness of the web? [Though isn't spam more of a problem for email at the moment?]
Andy Duncan, CEO of Channel 4 - web as 'tool of humanity', ability for humans to interact. Practical challenges to be solved. £50million 4IP fund. How do we get, grow ideas and bring them to the wider public, and realise the positive potential of ideas. Battle between positive public benefit vs economic or political aspects.
The internet brings more/different perspectives, but people are less open to new ideas - they get cosy, only talk to like-minded people in communities who agree with each other. How do you get people engaged in radical and positive thinking? [This is a really good observation/question. Does it have to do with the discoverability of other views around a topic? Have we lost the serendipity of stumbling across random content?]
Open to questions. 'Terms and conditions' - all comments must have a question mark at the end of them. [I wish all lectures had this rule!]
Questions from the floor: 1. why is the semantic web taking so long; 2. 3D web; 3. kids.
TBL on semantic web - lots of exponential growth. SW is more complicated to build than HTML system. Now has standard query language (SPARQL). Didn't realise at first that needed a generic browser and linked open data. (Moving towards real world).
[This is where I started to think about the question I asked, below - cultural heritage institutions have loads of data that could be open and linked, but it's not as if institutions will just let geeks like me release it without knowing where and why and how it will be used - and fair enough, but then we need good demonstrators. The idea that the semantic web needs lots of acronyms (OWL, GRDDL, RDF, SPARQL) in place to actually happen is a perception I encounter a lot, and I wanted an answer I could pass on. If it's 'straight from the horse's mouth', then even better...]
Questions from twitter (though the guy's laptop crashed): 4. will Google own the world? What would Channel 4 do about it?; 5. is there a contradiction between [collaborative?] open platform and spam?; 6. re: education, in era of mass collaboration, what's the role of expertise in a new world order? [Ooh, excellent question for museums! But more from the point of view of them wondering what happens to their authority, especially if their collections/knowledge start to appear outside their walls.]
AD: Google 'ferociously ambitious in terms of profit', fiercely competitive. They should give more back to the UK considering how much they take out. Qu to TBL re Google, TBL did not bite but said, 'tremendous success; Google used science, clustering algorithms, looked at the web as a system'.
CL re qu 5 - the web works best through norms and social interactions, not rules. Have to be careful with assumption that can regulate behaviour -> 'norm based behaviour'. [But how does that work with anti-social individuals?]
TBL re qu 6: e.g. MIT Courseware - experts put their teaching materials on the web. Different people have different levels of expertise [but how are those experts recognised in their expert context? Technology, norms/links, a mixture?]. More choice in how you connect - doesn't have to be local. Being an expert [sounds exhausting!] - connect, learn, disseminate - huge task.
Questions from the floor: 7. ISPs as villains, what can they do about it?; 9. why can't the web be designed to use existing social groups? [I think, I was still recovering from asking a question]
So the middle question was me. It should have been something like "if there's a tension between the top-down projects that don't work, and simple protocols like HTML that do, and if the requirements of the 'Semantic Web' are top-down (and hard), how do we get away from the idea that the semantic web is difficult to just have the semantic web?'" but it came out much more messily than that and the start of the answer was "Who told you SW is top down?". It was a leading question so it's my fault, but the answer was worth asking a possibly stupid/leading question.
TBL re qu 7 and ISPs 'give me non-discriminatory access and don't sell my clickstream'. [Hoorah!]
TBL re qu 8 - semantic web is bottom up, middle out, not top down. You can have lots of different data systems that talk different languages, imagine them as pieces of a quilt joined at the edges. World works like plug of stuff in sink, putting together lots of communities at lots of different levels. Take an existing set of terms unique to you, a small dataset, share it, negotiate with others to use those terms [or map between yours and theirs?], push to standards, use them on semantic web terms. Join interest groups. You can also do it from the top down. Global communities are hard work to make, lots of local communities are easy to make, the ones with important benefits are in the middle. The semantic web is the first technology designed with understanding that that's the way the world is. Scale free, 'fractal'.
[So I was asking 'how do we get to the semantic web' with reference to the museum sector - we can do this. Put a dataset out there, make connections to the organisation next to you (or get your users to by gathering enough anonymised data on how they link items through searching and browsing). Then make another connection, and another. We could work at the sector (national or international) level too (stable permanent global identifiers would be a good start) but start with the connections. "Small pieces loosely joined" -> "small ontologies, loosely joined". Can we make a manifesto from this?
There's also a good answer in this article, Sir Tim Talks Up Linked Open Data Movement on internetnews.com.
"He urged attendees to look over their data, take inventory of it, and decide on which of the things you'd most likely get some use out of re-using it on the Web. Decide priorities, and benefits of that data reuse, and look for existing ontologies on the Web on how to use it, he continued, referring to the term that describes a common lexicon for describing and tagging data."
Anyway, on with the show.]
Questions from the floor: 10. do we have enough "bosses who don't say no"?; 11. web to solve problems, social engineering [?]; 12. something on Rio meeting [didn't get it all].
TBL re 10 - he can't emulate other bosses but he tries to have very diverse teams, not clones of him/each other, committed, excited people and 'give them spare time to do things they're interested in'. So - give people spare time, and nurture the champions. They might be the people who seem a bit wacky [?] but nurture the ones who get it.
Qu 11 - conflicting demands and expectations of web. TBL - 'try not to think of it as a thing'. It's an infrastructure, connections between people, between us. So, are we asking too much of us, of humanity? Web is reflection of humanity, "don't expect too little".
TBL re qu 12 - internet governance is the Achilles heel of the web. No permission required except for domain name. A 'good way to make things happen slowly is to get a bureaucracy to govern it'. Slowness, stability. Domain names should last for centuries - persistence is a really important part of the web.
CL re qu 11 - possibilities of self-governance, we ask too little of the web. Vision of open, collaborative web capable of being used by people to solve shared problems.
JK - (NESTA) don't prescribe the outcome at the beginning, commitment to process of innovation.
Then Nesta hosted drinks, then we went to the pub and my lovely mate said "I can't believe you trolled Tim Berners-Lee". [I hope I didn't really!]
Monday, 23 June 2008
The BBC, accessibility, the hCalendar microformat and RDFa
Our concerns were:They're looking at using RDFa, which they describe as 'a slightly bigger S semantic web technology similar to microformats but without some of the more unexpected side-effects'.Until these issues are resolved the BBC semantic markup standards have been updated to prevent the use of non-human-readable text in abbreviations.
- the effect on blind users using screen readers with abbreviation expansion turned on where abbreviations designed for machines would be read out
- the effect on partially sighted users using screen readers where tool tips of abbreviations designed for machines would be read out
- the effect of incomprehensible tooltips on users with cognitive disabilities
- the potential fencing off of abbreviations to domains that need them
Their support for RDFa is timely in light of Lee Iverson's presentation at the UK Museums on the Web conference (my notes). It's also an interesting study of what can happen when geek enthusiasm meets existing real world users.
More generally, does the fact that an organisation as big as the BBC hasn't yet produced an API mean that creating an API is not a simple task, or that the organisational issues are bigger than the technical issues?
Thursday, 19 June 2008
Notes from 'UK Museums on the Web Conference 2008'
UK Museums on the Web Conference 2008 and the mashed museum day.
In the interests of getting my notes up quickly I'm putting them up pretty much 'as is', so they're still rough around the edges. I'll add links to the speaker slides when they are all online. Some photos from the two days are online - a general search for ukmw08 on Flickr will find some. I have some in a set online now, others are still to come, including some photos of slides so I'll update this as I check the text from the slides. These are my notes from the first session.
The keynote speech was given by Tom Loosemore of Ofcom on the Future of Public Service Content.
[For context, Ofcom is the 'independent regulator and competition authority for the UK communications industries' and their recently second review of public service broadcasting, 'The Digital Opportunity', caused a stir in the digital cultural heritage world for its assessment of the extent to which public sector websites delivered on 'public service purposes and characteristics'. You can read the summary or download the full report.]
'How many of you are on the main board of your institution?'
Leadership doesn't have the vision in place to take advantage of the internet.
Sees the internet as platform for public service, [most importantly] enlightenment. He's here today to enlist our help.
We view the internet through lens of expectations from the past, definitely in public service broadcasting - 'let's get our programs on the internet'.
What is value for money?
Would that other sectors did the same soul searching
[On the Ofcom review:] 'You can't really review the web, it's bonkers'
Public service characteristics to create a report card. Of the public service characteristics in the online market (high quality, original, innovative, challenging, engaging, discoverable and accessible), 'challenging' is the hardest.
Museums and cultural sector have amazing potential. What are the barriers between the people here who get it and being able to take that opportunity and redefine public service broadcasting?
It's not skills. Maybe ten years ago, not today. And it's not technology. The crucial missing link is leadership and vision, the lack of recognition by people who govern direction of institutions of the huge potential.
[Which does translate into 'more resources', eventually, but perhaps the missing gap right now is curatorial/interpretative resources? Every online project we do generates more enquiries, stretching these people further, and they don't have time to proactively create content for ad hoc projects as it is, especially as their time tends to be allocated a long time in advance.]
What's behind that reluctance, what can you do to help people on your board understand the opportunities? We can ask 'what business are we in? what's the purpose of our institution?'.
Tate recognise they're not just in the business of getting people to go to the Tate venues, they're in the business of informing people about art. Compare that to the Royal Shakespeare Company which is using its online site purely to get bums on seats.
Next opportunity... how do you take opportunity to digitise your collections and reach a whole new audience? How can you make better use of cultural objects that were previously constrained by physicalty.
What opportunities are native to the internet, can only happen there? How can it help your institution to deliver its purpose?
Recognise that you are in the (public service) media business.
How do you measure enlightenment? You could be changing the way people see the world, etc. but you need to measure it to make a case, to know whether you're succeeding. Metrics really really matter in public service arena.
BBC used to look at page views, but developers gamed the system. Then the metric was 'time online', but it stopped people thinking externally. Metric as proxy for quality.
Value = reach x quality. What kind of experience did they have?
Quality is the really hard part. As defined by BBC: quality is in the eye of the beholder. Did the user have an excellent experience?
BBC measure 'net promoter' - how likely are you to recommend this to a friend or colleague, on a scale of 1 - 10?
[But for our sector, what if you don't have any friends with the same interest in x? Would people extrapolate from their specific page on a Roman buckle to recommend the site generally?]
Throw away the 'soggy British middle' - the 7, 8s (out of ten).
Group them as Promoters (9-10/10), Passive (7-8/10), Detractors (0 - 6/10). The key measure is the difference between how many Promoters and how many Detractors. This was 'fabulously useful' at the BBC. 30% is good benchmark.
They mapped whole BBC portfolio against 'net promoters' % and reach, bubbles show cost.
It's not necessarily about reaching mass audiences. But when producing for niche audiences - they must love it, and it shouldn't cost that much.
He's telling us this because it's the language of funders, of KPIs, this is hard evidence with real people. You might use a different measure of quality but you can't talk about opportunities in abstract, must have numbers behind them.
Suggested the BBC's 15 Web Principles, including 'fall forward, fast'.
A measure of personal success for him would be that in x years when he asked 'who here is on the board of your institution, at least x should put hands up'.
[I really liked this keynote speech as a kick up the arse in case we started to get too complacent about having figured out what matters to us, as museum geeks. It doesn't count unless we can get through our organisations and get that content out to audiences in ways they can use (and re-use).]
In linking the sessions, Ross Parry mused about the legacy of 18th, 19th century ideas of how to build a museum, how would they be different if museums were created today?
Lee Iverson, How does the web connect content? "Semantic Pragmatics"
'Profoundly disagreed' with some of the things Tom was talking about, wants to have a dialogue.
He asked how many know the background to semantic web stuff? Quite a few hands were raised.
Talking about how the web works now and where it's going. Museums have significant opportunity to push things forward, but must understand possibilities and limitations.
Changing classic relationship - museum websites as face of institution to users. Huge opportunity for federating and aggregating content (between museums) - an order of magnitude better.
He's working with 13 museums, with north west native American artefacts. Communities are co-developers, virtually repatriating their (land).
Possibility to connect outside the museum. Powerhouse Museum as an excellent example of why (and how) you should connect.
Becoming connected:
Expose own data from behind presentation layers
Find other data
Integrate - creating a cohesive (situation)
Engage with users
Access to data is core business, curatorial stuff.
RDFa
Pragmatics of standards - get a sense of what it is you're doing [and start, don't try and create the system of everything first], it'll never work. Use existing standards if possible, grab chunks if you can. Never standardise what you minimally need to do to get the utility you need at the moment. Then extend, layers, version 2. A standard is an agreement between a minimum of two people [and doesn't have to be more complicated than that].
"Just do it" - make agreements, get it to work, then engage in the standardisation process.
Relationship between this and semantic web? Semantic web as 'data web'. Competing definitions.
Slide on Tim Berners-Lee on the semantic web in 1999.
Why hasn't it appeared? It's vapourware, you can't make effective standards for it.
Syntax - capability of being interpreted. Semantic - ability to interpret, and to connect interpretations.
Finding data - how much easier would it be if we could just grab the data we want directly from where we want it?
Key is relating what you're doing to what they're doing.
XML vs RDF
Semantic web built on RDF, it's designed for representing metadata. It's substantially different to XML. Lots of reaction against RDF has been reaction against XML encoding, syntactic resistance.
RDF is designed to be manipulated as data, XML is about annotating text. In XML, syntax is the thing, with RDF the data is the thing.
Grab entire XML doc before you can figure out how to smoosh then together. RDF works by reference, you can just build on it.
RDFa. A way of embedding RDF content directly in XHTML, relies on same strategies as microformats. Will be ignored by presentation oriented systems but readable by RDF parsers.
[RDF triples vs machine tags? RDF vs microformats? How RDF-like is OAI PMH?]
You can talk about things you don't have a representation for e.g. people.
Ignore the term 'ontology' - it's just a way of talking about a vocabulary.
Four steps for widespread adoption:
Promote practical applications
Develop applications now
[and the slide was gone and I missed the last two steps!]
There was also some stuff on limitations of lightweight approaches, and hermetically sealed museum data, user experiences. Also a bit on 'give away structured data' but with a good awareness of the need to keep some data private - object location and value, for example.
Ross - we've had the media context and technical context, now for the sector context.
Paul Marty, Engaging Audiences by connecting to collections online.
Vital connections...
What does it mean to say x% of your collection is online? For whom is it useful?
How to engage audiences around your collections? Not just presenting information.
Goes beyond providing access to data. Research shows audiences want engagement. Surveyed 1200 museum visitors about their requirements. [I would love to see the research] Virtuous circle between museum visits and website visits.
Build on interest, give experience that grabs people.
Romans in Sussex website - multiple museums offering collections for multiple audiences. Re-presenting same content in different ways on the fly.
Audiences
Don't just give general public a list of stuff. Give them a way to engage.
"Engaging a community around a collection is harder than providing access to data about a collection"
Photo of the week - says "What do you know about this photo? Please share your thoughts with us" But no link or instructions on how to do it. But at least they're trying...
Discussion - Tom, Lee and Paul.
"Why do you digitise collections before had need in mind?" [Because the driver is internal, not external, needs, would be the generous answer; because they could get funding to do it would be my ungenerous answer].
Tom on RDF - how seriously engaged with it to build audiences, tell stories.
BBC licence terms - couldn't re-use data for commercial purposes/at all.
Leadership need to understand opportunities because otherwise they won't support geek stuff.
Qu: terms of engagement - how is it defined?
Paul - US has made same mistakes re digitisation of collections and websites that don't have reusable data.
Participants must be involved in process from the beginning, need input at start from intended users on how it can engage them.
Fiona: why not use existing resources, go to existing sites with established audiences?
Lee: how did YouTube succeed - people were brought by embedded content. [This issue of using 'wrappers' around your content to help it go viral by being embeddable elsewhere was raised in another session too.]
Tom: letting go is how you win, but it's a profound challenge to institutions and their desire to maintain authority.
Friday, 6 June 2008
Yahoo! SearchMonkey, the semantic web - an example from last.fm
The first version of our application deals with artist, album and track pages giving you a useful extract of the biography, links to listen to the artist if we have them available, tags, similar artists and the best picture we can muster for the page in question.Some background on SearchMonkey from ReadWriteWeb:
At the same time, it was clear that enhancing search results and cross linking them to other pieces of information on the web is compelling and potentially disruptive. Yahoo! realized that in order to make this work, they need to incentivize and enable publishers to control search result presentation.(From Making the Web Searchable: The Story of SearchMonkey)
...
SearchMonkey is a system that motivates publishers to use semantic annotations, and is based on existing semantic standards and industry standard vocabularies. It provides tools for developers to create compelling applications that enhance search results. The main focus of these applications is on the end user experience - enhanced results contain what Yahoo! calls an "infobar" - a set of overlays to present additional information.
...
SearchMonkey's aim is to make information presentation more intelligent when it comes to search results by enabling the people who know each result best - the publishers - to define what should be presented and how.
And from Yahoo!'s search blog:
This new developer platform, which we're calling SearchMonkey, uses data web standards and structured data to enhance the functionality, appearance and usefulness of search results. Specifically, with SearchMonkey:This could be an interesting new development - the question is, how well does the data we currently output play with it; could we easily adapt our pages so they're compatible with SearchMonkey; should we invest the time it might take? Would a simple increase in the visibility and usefulness of search results be enough? Could there be a greater benefit in working towards federated searches across the cultural heritage sector or would this require a coordinated effort and agreement on data standards and structure?
- Site owners can build enhanced search results that will provide searchers with a more useful experience by including links, images and name-value pairs in the search results for their pages (likely resulting in an increase in traffic quantity and quality)
- Developers can build SearchMonkey apps that enhance search results, access Yahoo! Search's user base and help shape the next generation of search
- Users can customize their search experience with apps built by or for their favorite sites
Update to link to the Yahoo! Search Blog post ;The Yahoo! Search Gallery is Open for Business' which has a few more examples.
Monday, 26 May 2008
Notes from 'How Can Culture Really Connect? Semantic Front Line Report' at MW2008
The paper, "Semantic Dissonance: Do We Need (And Do We Understand) The Semantic Web?" (written by Ross Parry, Jon Pratty and Nick Poole) and the slides are online. The blog from the original Semantic Web Think Tank (SWTT) sessions is also public.
These notes are pretty rough so apologies for any mistakes; I hope they're a bit useful to people, even though it's so late after the event. I've tried to include most of what was discussed but it's taken me a while to catch up.
There's so much to see at MW I missed the start of this session; when we arrived Ross had the participants debating the meaning of terms like 'Web 2.0', 'Web 3.0', 'semantic web, 'Semantic Web'.
So what is the semantic web (sw) about? It's about intelligent and efficient searching; discovering resources (e.g. URIs of picture, news story, video, biographical detail, museum object) rather than pages; machine-to-machine linking and processing of data.
Discussion: how much/what level of discourse do we need to take to curators and other staff in museums?
me: we need to show people what it can do, not bother them with acronyms.
Libby Neville: believes in involving content/museum people, not sure viewing through the prism of technology.
[?]: decisions about where data lives have an effect.
Slide 39 shows various axes against which the Semantic Web (as formally defined) and the semantic web (the SW 'lite'?) can be assessed.
Discussion: Aaron: it's context-dependent.
'expectations increase in proportion to the work that can be done' so the work never decreases.
sw as 'webby way to link data'; 'machine processable web' saves getting hung up on semantics [slide 40 quoting Emma Tonkin in BECTA research report, ‘If it quacks like a duck…’ Developments in search technologies].
What should/must/could we (however defined) do/agree/build/try next (when)?
Discussion: Aaron: tagging, clusters. Machine tags (namespace: predicate: value).
me: let's build semantic webby things into what we're doing now to help facilitate the conversations and agreements, provide real world examples - attack the problem from the bottom up and the top down.
Slide 49 shows three possible modes: make collections machine-processable via the web; build ontologies and frameworks around added tags; develop more layered and localised meaning. [The data (the data around the data) gets smarter and richer as you move through those modes.]
I was reminded of this 'mash it' video during this session, because it does a good jargon-free job of explaining the benefits of semantic webby stuff. I also rather cynically tweeted that the semantic web will "probably happen out there while we talk about it".
Wednesday, 21 May 2008
BBC on microformats, abbr (and why the machine-readable web is good)
The web is a wonderful place for humans but it's a less friendly place for machines. When we read a web page we bring along our own learning, mental models and opinions. The combination of what we read and what we know brings meaning. Machines are less bright.
Given a typical TV schedule page we can easily understand that Eastenders is on at 7:30 on the 15th May 2008. But computers can't parse text the way we can. If we want machines to be able to understand the web (and there are many reasons we might want to) we have to be more explicit about our meaning.
Which is where microformats come in. They're a relatively new technology that allow publishers to add semantic meaning to web pages. These might be events, contact details, personal relationships, geographic locations etc. With this additional machine friendly data you can add events from a web page directly to your calendar, contacts to your address book etc. In theory it's a great combination of a web for people and a web for machines. But it has some potential problems.
One potential problem is microformat's use of something called the abbreviation design pattern.
Basically, if you have a screen reader and have abbreviation expansion turned on, they'd like to hear from you.
This overloading of the abbreviation tag also has implications for people using abbr correctly. It's a nice inline way to help explain jargon, but if browsers and screen readers change the way they parse and present the content, we'll lose that functionality.
The BBC guys also have a very interesting post on 'Helping machines play with programmes'.
Thursday, 15 May 2008
Notes from 'Aggregating Museum Data – Use Issues' at MW2008
These notes are pretty rough so apologies for any mistakes; I hope they're a bit useful to people, even though it's so late after the event. I've tried to include most of what was covered but it's taken me a while to catch up on some of my notes and recollection is fading. Any comments or corrections are welcome, and the comments in [square brackets] below are me. All the Museums and the Web conference papers and notes I've blogged have been tagged with 'MW2008'.
This session was introduced by David Bearman, and included two papers:
Exploring museum collections online: the quantitative method by Frankie Roberto and Uniting the shanty towns - data combining across multiple institutions by Seb Chan.
David Bearman: the intentionality of the production of data process is interesting i.e. the data Frankie and Seb used wasn't designed for integration.
Frankie Roberto, Exploring museum collections online: the quantitative method (slides)
He didn't give a crap of the quality of the data, it was all about numbers - get as much as possible to see what he could do with it.
The project wasn't entirely authorised or part of his daily routine. It came in part from debates after the museum mash-up day.
Three problems with mashing museum data: getting it, (getting the right) structure, (dealing with) dodgy data
Traditional solutions:
Getting it - APIs
Structure - metadata standards
Dodgy data - hard work (get curators to fix it)
But it doesn't have to be perfect, it just has to be "good enough". Or "assez bon" (and he hopes that translation is good enough).
Options for getting it - screen scrapers, or Freedom of Information (FOI) requests.
FOI request - simple set of fields in machine-readable format.
Structure - some logic in the mapping into simple format.
Dodgy data - go for 'good enough'.
Presenting objects online: existing model - doesn't give you a sense of the archive, the collection, as it's about the individual pages.
So what was he hoping for?
Who, what, where, when, how. ['Why' is the other traditional journalists questions but too difficult in structured information]
And what did he get?
Who: hoping for collection/curator - no data.
What: hoping for 'this is an x'. Instead got categories (based on museum internal structures).
Where: lots of variation - 1496 unique strings. The specificity of terms varies on geographic and historical dimensions.
When: lots of variation
How: hoping for donation/purchase/loan. Got a long list of varied stuff.
[There were lots of bits about whacking the data together that made people around me (and me, at times) wince. But it took me a while to realise it was a collection-level view, not an individual object view - I guess that's just a reflection of how I think about digital collections - so that doesn't matter as much as if you were reading actual object records. And I'm a bit daft cos the clue ('quantitative') was in the title.
A big part of the museum publication process is making crappy date and location and classification data correct, pretty and human-readable, so the variation Frankie found in data isn't surprising. Catalogues are designed for managing collections, not for publication (though might curators also over-state the case because they'd always rather everything was tidied than published in a possible incorrect or messy state?).
It would have been interesting to hear how the chosen fields related to the intended audience, but it might also have been just a reasonable place to start - somewhere 'good enough' - I'm sure Frankie will correct me if I'm wrong.]
It will be on museum-collections.org. Frankie showed some stuff with Google graph APIs.
Prior art - Pitt Rivers Museum - analysis of collections, 'a picture of Englishness'.
Lessons from politics: theyworkforyou for curators.
Issues: visualisations count all objects equally. e.g. lots of coins vs bigger objects. [Probably just as well no natural history collections then. Damn ants!]
Interactions - present user comments/data back to museums?
Whose role is it anyway, to analyse collections data? And what about private collections?
Sebastian Chan, Uniting the shanty towns - data combining across multiple institutions (slides)
[A paraphrase from the introduction: Seb's team are artists who are also nerds (?)]
Paper is about dealing with the reality of mixing data.
Mess is good, but... mess makes smooshing things together hard. Trying to agree on standards takes a long time, you'll never get anything built.
Combination of methods - scraping + trust-o-meter to mediate 'risk' of taking in data from multiple sources.
Semantic web in practice - dbpedia.
Open Calais - bought out from Clearforest by Reuters. Dynamically generated metadata tags about 'entities' e.g. possible authority records. There are problems with automatically generated data e.g. guesses at people, organisations, whatever might not be right. 'But it's good enough'. Can then build onto it so users can browse by people then link to other sites with more information records about them in other datasets.
[But can museums generally cope with 'good enough'? What does that do to ideas of 'authority'? If it's machine-generated because there's not enough time for a person in the museum to do it, is there enough time for a person in the museum to clean it? OTOH, the Powerhouse model shows you can crowdsource the cleaning of tags so why not entities. And imagine if we could connect Powerhouse objects in Sydney with data about locations or people in London held at the Museum of London - authority versus utility?
Do we need to critically examine and change the environment in which catalogue data is viewed so that the reputation of our curators/finds specialists in some of the more critical (bitchy) or competitive fields isn't affected by this kind of exposure? I know it's a problem in archaeology too.]
They've published an OpenSearch feed as GeoRSS.
Fire eagle, Yahoo beta product. Link it to other data sets so you can see what's near you. [If you can get on the beta.]
I think that was the end, and the next bits were questions and discussion.
David Bearman: regarding linked authority files... if we wait until everything is perfect before getting it out there, then "all curators have to die before we can put anything on the web", "just bloody experiment".
Nate (Walker): is 'good enough' good enough? What about involving museums in creating better and correcting data. [I think, correct me if not]
Seb: no reason why a museum community shouldn't create an OpenCalais equivalent. David: Calais knows what reuters know about data. [So we should get together as a sector, nationally or internationally, or as art, science, history museums, and teach it about museum data.]
David - almost saying 'make the uncertainty an opportunity' in museum data - open it up to the public as you may find the answers. Crowdsource the data quality processes in cataloguing! "we find out more by admitting we know less".
Seb - geo-location is critical to allowing communities to engage with this material.
Frankie - doing a big database dump every few months could be enough of an API.
Location sensitive devices are going to be huge.
Seb - we think of search in a very particular way, but we don't know how people want to search i.e. what they want to search for, how they find stuff. [This is one of the sessions that made me think about faceted browsing.]
"Selling a virtual museum to a director is easier than saying 'put all our stuff there and let people take it'".
Tim Hart (Museum Victoria) - is the data from the public going back into the collection management system? Seb - yep. There's no field in EMu for some of the stuff that OpenCalais has, but the use of it from OpenCalais makes a really good business case for putting it into EMu.
Seb - we need tools to create metadata for us, we don't and won't have resources to do it with humans.
Seb - Commons on Flickr is good experiment in giving stuff away. Freebase - not sure if go to that level.
Overall, this was a great session - lots of ideas for small and large things museums can do with digital collections, and it generated lots of interesting and engaged discussion.
[It's interesting, we opened up the dataset from Çatalhöyük for download so that people could make their own interpretations and/or remix the data, but we never got around to implementing interfaces so people could contribute or upload the knowledge they created back to the project, or how to use the queries they'd run.]
Saturday, 19 April 2008
Explaining the semantic web: by analogy and by example
- Web 1.0 is like buying a can of Campbell's Soup
- Web 2.0 is like making homemade soup and inviting your soup-loving friends over
- The semantic web is like having a dinner party, knowing that Tom is allergic to gluten, Sally is away til next Thursday and Bob is vegetarian.
To extend the analogy, it's also as if the semantic web could understand that when your American aunt's soup recipe says 'cilantro', you'd look for 'coriander' in shops in Australia or the UK.
Explaining by doing: this review 'Why I Migrated Over to Twine (And Other Social Services Bit the Dust)' of Twine gives lots of great examples of how semantic web stuff can help us:
So for example when Stanley Kubrick is mentioned in the bookmarklet fields, or in the document you upload, or in the email you send into Twine — the system will analyze and identify him as a person (not as a mere keyword). This is called entity extraction and is applied to all text on Twine.
Under the hood, a person is defined in a larger ontology in relation to other things. Here’s an example of a very small portion of my own graph within Twine:
Some may not find the point of this clear. So to explain: Just as HTML enables computers to display data — this extra semantic information markup (RDF, OWL, etc.) enables computers to understand what the data is they’re displaying. And moreover, to understand what things are in relation to other things.
Example Search
For an example, when we search for “Stanley Kubrick” on regular search engines, the words “Stanley” and “Kubrick” are usually regarded as mere keywords: a series of letters that the search engine then tries to find pages with those series of letters. But in the world of semantic web, the engines know “Stanley Kubrick” is a person. This results in a lot less irrelevant items from the search’s results....
If you weren’t already aware, the systems I just described above are the basic semantic web concept: Encapsulating data in a new layer of machine processable information to help us search, find and organize the overwhelming and ever-growing sea of pictures, videos, text and whatever else we’re creating.
I think these are both useful when explaining the benefits of the semantic web to non-geeks and may help overcome some of the fear of the unknown (or fear of investment in the pointless buzzword) we might encounter. If we believe in the semantic web, it's up to us to explain it properly to other people it's going to effect.
I also discovered a good post by Mike on the 'Innovation Manifesto'.
Friday, 18 April 2008
It's a wonderful, wonderful web
This experiment is part of Google's broader effort to increase its coverage of the web. In fact, HTML forms have long been thought to be the gateway to large volumes of data beyond the normal scope of search engines. The terms Deep Web, Hidden Web, or Invisible Web have been used collectively to refer to such content that has so far been invisible to search engine users. By crawling using HTML forms (and abiding by robots.txt), we are able to lead search engine users to documents that would otherwise not be easily found in search engines, and provide webmasters and users alike with a better and more comprehensive search experience.You're probably already well indexed if you have a browsable interface that leads to every single one of your collection records and images and whatever; but if you've got any content that was hidden behind a search form (and I know we have some in older sites), this could give it much greater visibility.
Secondly, Mike Ellis has done a sterling job synthesising some of the official, backchannel and informal conversations about the semantic web at MW2008 and adding his own perspective on his blog.
Talking about Flickr's 20 gazillion tags:
So far, so ace. We've been excited about using the implicit links created between data as people consciously record information with tags, or unconsciously with their paths between data to create those 'small ontologies, loosely joined'; the possibilities of multilingual tagging, etc, before. Tags are cool.To take an example: at the individual tag level, the flaws of misspellings and inaccuracies are annoying and troublesome, but at a meta level these inaccuracies are ironed out; flattened by sheer mass: a kind of bell-curve peak of correctness. At the same time, inferences can be drawn from the connections and proximity of tags. If the word “cat” appears consistently - in millions and millions of data items - next to the word “kitten” then the system can start to make some assumptions about the related meaning of those words. Out of the apparent chaos of the folksonomy - the lack of formal vocabulary, the anti-taxonomy - comes a higher-level order. Seb put it the other way round by talking about the “shanty towns” of museum data: “examine order and you see chaos”.
The total “value” of the data, in other words, really is way, way greater than the sum of the parts.
But the applications of this could go further:
I got thinking about how this can all be applied to the Semantic Web. It increasingly strikes me that the distributed nature of the machine processable, API-accessible web carries many similar hallmarks. Each of those distributed systems - the Yahoo! Content Analysis API, the Google postcode lookup, Open Calais - are essentially dumb systems. But hook them together; start to patch the entire thing into a distributed framework, and things take on an entirely different complexion....
Here’s what I’m starting to gnaw at: maybe it’s here. Maybe if it quacks like a duck, walks like a duck (as per the recent Becta report by Emma Tonkin at UKOLN) then it really is a duck. Maybe the machine-processable web that we see in mashups, API’s, RSS, microformats - the so-called “lightweight” stuff that I’m forever writing about - maybe that’s all we need. Like the widely accepted notion of scale and we-ness in the social and tagged web, perhaps these dumb synapses when put together are enough to give us the collective intelligence - the Semantic Web - that we have talked and written about for so long.I'd say those capital letters in 'Semantic Web' might scare some of the hardcore SW crowd, but that's ok, isn't it? Semantics (sorry) aside, we're all working towards the same goal - the machine-processable web.
And in the meantime, if we can put our data out there so others can tag it, and so that we're exposing our internal 'tags' (even if they have fancier names in our collections management systems), we're moving in the right direction.
(Now I've got Black's "Wonderful Life" stuck in my head, doh. Luckily it's the cover version without the cheesy synths).
Right, now I'm off to the Museum in Docklands to talk about MultiMimsy database extractions and repositories. Rock.
