Wednesday, August 31, 2011

Ideas on Linked Open Transportation Data for TravelHack

Earlier this summer I saw some tweets about a nice event here in Gothenburg: West Coast TravelHack 2011, 8-9 October. As I am a daily commuter (with Västtrafik's trams, buses and trains) and an information architect addicted to the linked data idea, and I also have a background as researcher in mobile informatics, I got two ideas and wrote them up as tweets (tweet 1 and tweet 2)

Today, I saw some tweets linking to two articles about the interesting FixMyTransportation:
Looking for hackers
I was reminded of my two ideas and also of my time as a part-time industrial PhD researcher. My research in the Mobile Informatics group, at the Victoria Institue and IT University, concerned the mechanisms needed to provide highly mobile professionals, such as new journalists, with contextualized information using mobile applications: "Mobile Newsmaking" (thesis, presentation)

So, I posted a tweet about FixMyTransportation it in Swedish and Karl-Petter Åkesson (@kallep), an old friend from my time as part-time researcher, kindly replied and said in his tweets back (tweet 1 and tweet 2): Why not get together with a couple of hackers and show how your ideas for a linked data infrastructure could enable nice apps and services for commuters. Great, I tweeted back -- but, I don't know that many great hackers as it's ten years since I did my research on mobile applications.

So, now I am looking for some great hackers to potentially explore my ideas on Linked Open Transportation Data at the TravelHack event 8-9 October.


Give every bus stop, tram route and train station etc. a URI
Identify "things" globally by using http based URIs (Uniform Resource Identifiers) - today all public schools, roads, ministers, and many bus stops, in UK have URIs

For example the URI http://transport.data.gov.uk/id/stop-point/1800SJH1081 identifies a bus stop in Manchester. Assigning a http based URI is what the two first principles of Linked Data say. 

The third principle say that you should provide useful information about the thing when its URI is dereferenced, using standard formats such as RDF/XML. So, if you you put this URI http://transport.data.gov.uk/id/stop-point/1800SJH1081  in a web browser it will give you a nice html documentation of the metadata describing the busstop. A app or service could choose between for example a RDF/XML file or a JSON file.  See my Linked Data page for some nice videos, books, blogs etc.

Use a common vocabulary for transportation
And all these "things" can also be typed, described and linked using classes, properties and relationships from a range of vocabularies for different domains. 

For the transportation domain I have seen some nice tweets pointing me to TRANSIT: A vocabulary for describing transit systems and routes

There are also many general vocabularies and ontologies that are commonly used to publish linked data. You can 'cherry-pick' from some of the most common, for example Friend-of-a-Friend (FOAF) provides terms for describing people and their social network, SIOC Semantically-Interlinked Online Communities, and Dublin Core defines general metadata attributes.

Kudos to @peterkz_swe, @egonwillighagen, @wieselgren, @kallep
for nice interactions on Twitter inspiring me to write this blog post

Thursday, August 18, 2011

A prediction 3-5 years from now

Making predictions can be tricky. However, a former colleague, and actually also my manager for a short time, Jean-Peter Fendrich (@carokanns) recently published a few predictions 3-5 years from now in the LinkedIn group Volvo IT Innovation Centre
  • Inspired by iPad and its competitors there will come a new device that replaces the Laptop as we know it now. 
  • Html5 will make all these app's and app technologies obsolete. 
  • We will finally have standards and infrastructure that support "mobile wallet" - replacing cash, credit cards and other payment systems.

JP asked for feedback and more predictions, so I posted the following:

As JP and I, together with Martin Börjesson (@futuramb), Annika Eriksson, Christian Forsäng  and Else-Marie (Emma) Malmek, were some of the folks introducing the first generation of web technology (Web 1.0) in the Volvo organisation back in the mid 90ies it was nice to highlight the third generation (Web 3.0) in this Volvo IT group.

The focus in my blog postings and tweets the last year or so has been on two of the fundaments for Web 3.0, i.e. the Linked Data principles and in particular the use of http based URIs. For more details, see one of my first blog posts: Corporate Transparency and Linked Data. See also my list on URI Design that I try to keep updated.

"Data is the new electricity. URIs are the conduction mechanism."
Quote by Kingsley Uyi Idehen (@kidehen)
  

Tuesday, August 9, 2011

ICBO2011 Reports

The last week in July I and three colleagues attended the International Conference on Biomedical Ontology (ICBO) 2011, in Buffalo, NY. As I have been a "remote hang-around" on Twitter following other conferences on distance (see for example my blog post following the SemTech conference earlier this summer) it was great fun this time to be active on Twitter IRL in Buffalo:  My #ICBO2011 tweets


And yes, I did see the Niagara Falls again -- this time I did get really close to them on a boat tour with the "Maid of the Mist".


Now, after a long journey home, and a couple of relaxing days on the Swedish west coast and in central London, it's time to use my tweets, the conference presentations and proceedings (pdf) to pull together some of my insights and learnings. Here's my first report with some notes and reflections from the conference and follow up to my previous blog posts in preparation for the conference (part 1 and part 2). See also my fourth blog post from ICBO published 1 September.


High quality, "true", ontologies 
It was nice to see presentations and read papers on ontologies from a broad spectrum of domains, such as:
  • Genes
    See a recent paper: How the Gene Ontology Evolves, describing the ways in which curators of the Gene Ontology (GO) have incorporated new knowledge. 
  • Protein complex and supra-complex
    See the presentation on this topic in the panel the first day: From proteins to diseases, by Bill Crosby (Department of Biological Sciences, University of Windsor)
  • Emotions and Chronic pain
    See the presentation and paper on how to represent emotions based on research in affective disorders such as bipolar, depression and schizoaffective disorder, by Janna Hastings, (European Bioinformatics Institute, UK, and, Swiss Centre for Affective Sciences, University of Geneva, Switzerland). See also the announcement of the development of an ontology for Chronic pain and a nice video: Toward a New Vocabulary of Pain.
  • Demographics
    See the presentation describing how "demographic data in current information systems is ad hoc, and current standards are insufficient to support accurate capture and exchange of demographic data", and the proposed use of the Demographics Application Ontology to as a solution. 
  • Adverse Events
    In the workshop on representing adverse events we learned about interesting work on adverse ontologies. (See a video of the workshop organizer Mélanie Courtot: Towards an Adverse Event Reporting Ontology). We also learned about the development of ontologies to represent temporal relationships (e.g. Clinical Narrative Temporal Relation Ontology) which is a key aspect in handling safety issues and regular ongoing pharmacovigilance in pharmaceutical research and development.
All of these are examples of high quality "true"1) and modular ontologies developed beneath the Basic Formal Ontology (BFO) providing formal definitions for types of entities in reality and for the relationships between such entities (so called ontological realism). Such ontologies are designed to allow annotations of experimental and clinical data "to be unified through  disambiguation of the terms employed in a way  that allows complex statistical and other  analyses to be performed which lead to the  computational discovery of novel insights"2)


My own reflections: 
So far we have seen none, or very little, uptake of such high quality "true" ontologies for clinical data. Something I also highlighted in my earlier blog post on clinical data standards.  In a coming blog post I will present a demo using the Demographics Application Ontology showing how a high quality "true" ontology can be used to support accurate capture and exchange of demographic data. I will also outline some ideas on how this could be used also for clinical study data (CRF:s and databases). 

"Mapping mania" for the legacy of terminologies
A common theme in several of the presentations, papers and panels was the mappings (matching, alignment) needed between terms and concepts organized as terminologies and coding nomenclatures, such as SNOMED CT, LOINC, ICD, CDISC SDTM CT:s (derived from NCI Thesaurus), and MedDRA. Here are some examples:
  • Extraction of the anatomy value set from SNOMED CT to be reused for the 11th revision of the International Classification of Diseases (ICD-11). See a presentation on the problems and proposed patterns by some well known people (Harold Solbrig and Christopher Chute at Mayo Clinic, Kent Spackman working for IHTSDO, and Alan L. Rector at University of Manchester)
  • The Ontology Evaluation Alignment Initiative (OAEI) was mentioned by several presenters as a forum to discuss the problems of direct matching between different terminological resources.
  • The use of a ontology matching tool called AgreementMaker was presented.
  • In a panel on: National Center for Biomedical Ontology (NCBO) Technology in Support of Clinical and Translational Science, the basic lexical term mappings was mentioned as an example of a service available both via BioPortal's graphical interface and as REST services.
These are all example of a legacy already in use, or in the process of being used, for the annotations of EHR, clinical trials and patient safety data. For example for the huge US initiative on meaningful use of EHR as highlighted by Roberto Roch in his keynote on Practical Applications of Ontologies in Clinical Systems.


My own reflections:
In my previous blog post preparing for the conference I refereed to the mapping problem as  comparing "Apples and Oranges" and sometimes I think of it as a "mapping mania". In the conference I did hear the comment "Mappings are hard" several times,  and also the question "Who will create, validate and maintain all the mappings?


After some more days of vacation I will get back later on in August with more notes and reflections from the conference:.
  • I will report from the debate on how to accurately connect data from measurements and questionnaires (information entities) to ontologies (real world entities). I think this is a key aspect to get machine-processable clinical data ready for automatic transformation and direct querying, and ready for inferencing and reasoning. 
  • Another theme I would like to cover is referent tracking, i.e. assign globally unique identifiers for each entity in reality about which information is stored. For example diagnoses, procedures, demographics, encounters, hypersensitivity, and observations as they are reported in EHRs. This is something I think is a key enabler for accurate secondary use of EHRs.

Friday, July 22, 2011

ICBO2011 Preparations, part two

Via the email lists for the Clinical Data Interchange Consortium (CDISC) Terminology team, and for the Electronic Health Records for Clinical Research (EHR4CR) one of the Innovative Medicines Initiative (IMI) project, I have see some recent discussions on cross-terminology mapping challenges. Challenges  due to the fact that terminologies and coding nomenclatures, such as SNOMED CT, LOINC, CDISC SDTM CT:s, and MedDRA, all have been developed for different purposes, with disparate approaches and structures.

Together with attendances from NCI, NCBO, FDA, Mayo, SAS, Stanford and other organizations, I and a few colleagues, will attend the International Conference on Biomedical Ontology (ICBO) next week . See my previous blogpost with some more background. 


Photo (Flickr): Automania

Apples and Oranges
In preparations for the workshop the first day, Representing Adverse Events, I did find this paper highly interesting as it compare and contrast SNOMED CT and MedDRA, and also describes the challenges in mapping between them: Heterogeneous but “standard” coding systems for adverse events: Issues in achieving interoperability between apples and oranges.



OBO Foundry based ontologies as "catalyst"
I hope the adverse event workshop, and the whole ICBO event, will be an opportunity for me to learn more about the Open Biology and Biomedical Ontologies (OBO) Foundry approach, and to discuss the challenges and opportunities in a “common language with which to energize cross-disciplinary research1)


I hope to better understand “how legacy terminologies, such as SNOMED CT, and the data coded with their aid can be successfully used for information-driven clinical and translational research2). My understanding is that the approach to be discussed at this event is the use of OBO Foundry based high-level reference ontologies, such as the Ontology for General Medical Science (OGMS), as a kind of catalyst instead of direct terminology-to-terminology mappings.


Yet another "standard", or ...
At the same time I did find this cartoon, circulating on Twitter this week, quite amusing. So, I think it will be a hot and interesting week in Buffalo, NY..
xkcd: Standards
Here's a brief introduction to the use case I and a colleague will present at the adverse event workshop:
"A use case will be presented describing how a query from a regulatory authority is handled as part of the regular ongoing pharmacovigilance in pharmaceutical research and development. It will illustrate how databases and literature are being reviewed manually, exemplify how different databases are structured and highlight some of issues in the coding of data. With this use case, we hope to provide a background to our interest in an ontologically based approach to enable a more automatic way to access, structure and analyze patient safety related data."

Tuesday, June 28, 2011

ICBO2011 Preparations

In a couple of weeks I will attend the International Conference on Biomedical Ontology (ICBO) 2011, in Buffalo, NY. 
In July, hundreds of international scientists from dozens of biomedical fields will meet at the University at Buffalo seeking a common language with which to energize cross-disciplinary research.“ From ICBO News: For the Sake of Research and Patient Care, Scientists Must Find Common Language
And yes, it will be a great opportunity for me to see the Niagara Falls again. This time  from the American side. Last time I saw it was in 1999 from the Canadian side when I attended the W3C conference in Toronto. The WWW8 conference where I was absolutely thrilled by the power of the simple and elegant model of RDF triples. At the WWW8 I also heard Tim Berners-Lee talk about the Semantic Web for the first time.

The coming weeks I hope to able to do a re-cap of a couple of ontology related papers and articles, and also read and digest some new ones listed for the events I have signed up for:

I will use one or two forthcoming blog posts to write up my insights and reflections coming to my mind while reading.

Here's a quote I think well captures my motivation to learn more about ontologies and getting my ICBO2011 attendance approved by my managers. It's taken from this great article More than Words: Biomedical Ontologies with references to the work of several of the international scientists who will get together at the ICBO2011.
“… true ontologies are more than just controlled terms. They capture, in a logical, systematic way, what scientists regard as the basic truths about a topic. Like equations in physics or axioms in mathematics, they can even be the basis for computational models. When connected to databases, scientific papers, and software applications, ontologies ‘help cope with the ever-growing, chaotic accumulation of text and facts" in biomedical and translational research.“

Sunday, June 12, 2011

SemTech2011

The last couple of days the Twitter feeds for #semanticweb and #linkeddata have been very busy and #semtech peaked with more than one tweet per minute during the Semantic Technology Conference 2011 in San Fransisco 5-9 June.
See the #SemTech 2011 Twitterscript for agreat overview of all the #semtech tweets sent during the conference, aligned with the sessions going on at the time. Kudos to  @glenn_mcdonald and @needlebase.
For me, here over in Sweden,  it's been a couple of late evenings and some busy mornings catching up on Twitter while commuting. Below some of the presentations, discussions, and blogs I did find extra interesting.

schema.org
A couple of days before the conference the news came out on Twitter about the announcement from Google, Yahoo and Microsoft (Bing) on their joint schema.org. A global, single vocabulary and the use of Microdata to encode structured data into web-pages using this vocabulary for search engines to do a better job.
A graph centric visualization of the schema.org vocabulary
with "Thing" in the center of it


The first comment I re-tweeted as a "I Liked" on this topic was a tweet on Friday 5 June by Darin L. Stewart (@darinlstewart) pointing to his posting on Gartner's blog: Schema.org: Webmaster One-Stop or Linked Data Land Grab? With some early critique. At the same time came the first version of a RDF Schema version of the vocabulary on schema.rdfs.org. Great job done by Michael Hausenblas (@MHausenblas) et al.. And I did find it interesting to read the quick, positive comment from Chris Bizer, the Linked Data guru behind DBpedia, on Google's official webmaster blog. During the conference schema.org was also the *hot* topic and late Wednesday evening my time I followed a heated online IRC discussion from the BOF on structured data in HTML and vocabularies. For more reading on this topic see the link bundle called schema.org is in town compiled by Michael Hasheke (@hashek)

Linked Data Tutorial and Cookbook
Among all the tutorials and presentations at the conference I picked up two great Linked Data resources, First of all Juan Sequeda's (@juansequeda) tutorial series, and also  a presentation "I liked, very much"- The Joy of Data - A cookbook for publishing and consuming Linked Data by Bernadette Hyland (@BernHylland). These two triggered me to create a separate Linked Data Resource Page with my favorites, including these two.

Linked Health Data
The last day of the conference I spotted some tweets that toke me to the presentation I liked most of all: Clinical quality linked data on health.data.gov, presented by George Thomas (@georgethomas). See also his blog post on data.gov with an excellent argumentation for linking  publicly available health data such as hospital compare data:
In addition to making flatfiles available to download on the Web, and providing applications that enable programmatic access to backend databases through the Web, imagine using the Web itself as a database: a massively distributed, decentralized database. This is what Linked Data is about – putting data in the Web.

Two technologies to catch up with
Many tweets talked two Calimachus, a framework for data-driven applications based on Linked Data principles allowing Web authors to quickly and easily create semantically-enabled Web applications. I will have a look at the Calimaschus videos they published. And a presentation on Semantic Architecture & Composing Resource Oriented System, by Brian Sletten (@bsletten), made me curios to learn more about the architecture thinking called  NetKernel.

Other blog posts 
I look forward to read several reflective blog posts the coming week when the participants are back home. For example, I look forward to see what Darin L. Stewart (@darinlstewart) will report from SemTech 2011 on his Gartner blog. I will update this blog post with links to what I find interesting.  

Monday, May 2, 2011

Linking Clinical Data Standards

This is a follow-up to an earlier blog post where I outlined the background, audiance and intention of three presentations. Two of them have been published on Slideshare:
Here I focus on the second presentation, a presentation I did in the CDISC (Clinical Data Interchange Standards Consortium) conference in Brussels recently. One of the key people in the CDISC community, Dave Iberson-Hurst, lists semantic web as one of three themes and kindly refers to my presentation in a recent blog post

My presentation, and also a very nice presentation from Roche, triggered interesting questions. Questions both on what I proposed as pragmatic first steps for linking clinical data standards, and also on what I see as future opportunities. Below you find the questions and my "answers", or rather thoughts. In a coming blog post I will discuss what all of this could mean for CDISC SHARE (metadata repository).

In my presentation - the last one on the first day - I  urged the CDISC community to consider the use of semantic web standards and linked data principles for clinical data standards. It was very nice to be able to refer back to two of the presentations in the earlier sessions. 

Pragmatic steps for CDISC
Firstly, to the presentation by Rebecca Kush, President of CDISC, on the value of open and free standards. The key message in my presentation pointed out:




Roche use Semantic Web for clinical data standards
And secondly, to the presentation from Roche on the development of a "Global Data Standard Repository" (GDSR) using semantic web standards and a ontology tool (TopBraid Composer). My first slides introducing the idea of "Triples" (the RDF standard model) and "Global Identifiers" (URI:s) was a recap for the audience as Frederik Malfait (IMOS Consulting presenting on behalf of Roche) in a really good way already had introduced these. 

Questions and Answers
Even though it was the last presentation for the day (just before  a very nice evening with TinTin at the Brussels Comic Strip Center) many people stayed around and I got the opportunity to sort out a key question, and also to outline two future opportunities: 
Q: Do you mean we should publish the actual clinical data openly? 
A: No! What should be made publicly available is another topic. My key message is that the free and open clinical data standards as they are currently constructed should be made available as linked open clinical data standards 1]. This means, using semantic web standards. (I propose the use of RDF/XML format as an alternative to Excel and ODM/XML.) And, also applying the Linked Data principles. (For example, assigning URI:s as global identifiers as an alternative to text strings for the submission values.)
Q: Does this relates to ontologies for bioinformatics?
A:
Yes. The insights from developing for example the Gene Ontology are highly applicable when representing and structuring the entities and relations in the clinical reality. In some extra slides to my presentation I propose explorative work to construct the next generation of clinical data standards using modern ontologies 2] based on the so called Open Biological and Biomedical Ontologies (OBO) Foundry.
Q: Do you mean that this would take away the need for manual transformation of clinical data?
A:
Yes and No.
Yes, because the above outlined next generation of clinical data standards (i.e. using semantic web standards, applying linked data principles and being based on modern ontologies) would improve the research utility of clinical datasets. That is, firstly, a very normalized, flexible way to convey clinical data. And, secondly, machine-processable clinical data ready for automatic transformation and direct querying, and ready for inferencing and reasoning.
No, because existing data needs to be transformed according to the above. And, No for quite some time as there are many things to explore and learn. A  highly pragmatic, incremental and stepwise approach is required 3] 

1]  
See my presentation slide 31-36 for more details on the pragmatic steps I propose for CDISC, and NCI.
2]  The two OBO Foundry based ontologies I am referring to are the Translational Medicine Ontology, TMO (a.k.a. the Pharma Ontology) and the Computer-Based Patient Record (CPR) Ontology. See also an excellent article on biomedical ontologies: More Than Words, in the Clinical and Translational Science Network.
 

Kudos to Frederik Malfait and Jonathan Chainey (Roche), 
Dave Iberson-Hurst (@Assero_UK),
Bron Kisler (@CDISC), Philippe Verplancke and
Isabelle de Zegher 
 for great discussions F2F in Brussels.