Monday, June 24, 2013

Semantic Trilogy preparation

The Swedish Midsummer weekend is over and it's time to look forward. Saturday 6th to Friday 12th of July I'll attend the Semantic Trilogy in Montreal, Qc, Canada.

I plan to attend these events during the week:
In 2011 I, together with three colleagues, attended the ICBO 2011 event (see my three blog post: Preparations part 1 and part 2,  report). So, I look forward to reconnect with people in the OBO (The Open Biological and Biomedical Ontologies) community.

And to meet F2F interesting people in the W3C HCLS (Semantic Web Health Care and Life Sciences Interest Group). And people interested in ontologies and semantic web working for e.g. Sanofi, Novo Nordisk, Mayo Clinic.

I'm also very happy that I'll get the opportunity to attend my third semantic web related event in Canada.
  • In 2007 I attended the WWW2007 conference in wonderful Banff.
"During the WWW2007 conference a breakthrough of the Linked Data idea happened in a session where web experts demonstrated the power of a new generation of the web, a web of data. For us attending the session it was hard to imagine the full potential on what this idea would mean for individual scientists and for a pharmaceutical company." 
From  Linked Data, an opportunity to mitigate complexity in pharmaceutical research and development, Bo Andersson and Kerstin Forsberg, LWDM 2011 

And yes, I do hope to also get some time during the weekend to visit the Jazz Festival.

Tuesday, June 11, 2013

Standards for common aspects

Through the last three years I have been engage with different groups working on standards, both for data exchange, such as CDISC, and for vocabularies such as MedDRA MSSO and NCI EVS. As they now start to see the value of using "standards for standards".


Push Back
From Flickr bitpuddle
(Twitter
@eric_d_hancock)

Standards for standards

So, "I push back" to standard organisations to use semantic web standards and linked data principles to make their standards directly usable for humans and for machines.

A good example is CDISC and their growing interest in using semantic web standards (based on RDF, Resource Description Framework): CDISC2RDF. For some background see Clinical studies and the road to Linked Data. Today FDA, CDISC, pharma:s, CRO:s and software vendors are working together on this in a FDA working group for Semantic Technology organised by PhUSE.



Standards for common aspects

The last year or so, I have also tried to keep up to date with groups developing RDF-based standards for common aspect such as:
  • data descriptions (VoID)
  • data provenance and versioning (PROV and PAV)
  • concept based vocabularies and value sets (SKOS)
  • multi-dimensional statistical data (RDF Data Cube)
I try to ensure that we have a good view of the maturity and applicability of these standars so we can use them in our internal“integration factory”. But most of all “push back” to vendors. I foresee that we in the same way started to add requirements on web-interfaces for better end user usability back in the late 90:ies, we now should start to add requirements on web-interfaces for better machne usability. So we need to to understand how to incorporate these common aspects in our URS:s, RFI:s RFP:s etc..

For software vendors to use RDF-based standards for common aspects, for example:
  • MediData's Rave and Perceptive's IMPACT to describe datasets using VoID.
  • Accelrys' Pipeline Pilot to use W3C PROV.
  • Microsoft's SharePoint to use term sets for tagging in SKOS.
  • SAS Institute's Drug Development to create analysis results using RDF Data Cube.

So, this interview with Reza B'Far, Vice President of Development, Oracle on the W3C blog made me vryy glad: Oracle on Data on the Web
Oracle to use W3C provenance standard to create a single audit time line across systems
"One of the hugest problems we faced was maintaining transaction audit trails in a heterogeneous environment in a standard and compatible way. Audit trails are described with literally millions of different formats in different organizations. This used to mean it was impossible to create a single audit time line. PROV solves this problem. We now provide (and consume) a PROV feed that unifies the audit trails generated by transactions across heterogeneous systems."
See also the Implementation report with 60+ examples of usage of the W3C Provenance specifications.

For a nice intro to the W3C Provenance Specifications, see the tutorial by Paul Groth (@pgroth) at the Extended (European) Semantic Web conference.


Saturday, May 25, 2013

Three Linked Data meetings in Sweden

I'm back after two nice day in the south of Sweden. Yesterday, 24 May, I attended the first meetup for Linked Data in Malmö.


This was the third Linked Data meeting in Sweden. They have all been great events with more than 30 attendees each. I do hope these will encourage more friends and colleagues In Sweden across academia, industry, consult companies and government to start applying the Linked Data principles and use the stack of Semantic Web standards. 

Links to all three events:
Kudos to Bosse Andersson (@bBalsa), Marie Gustavsson-Friberg (@mariegus)
and Eva Blomqvist (@evabl444) for arranging. I look forward the next one!

Sunday, March 31, 2013

Talking to machines

The last week I remotely followed two events while commuting, two events related to Evidence Based Medicine (EBM), both took place in Oxford:

+Ben Goldacre did speak at both events. At the Cochrane event he talked about getting better in talking to the Public, to Policy makers and to Machines. In the last part of his talk: Talking to Machines he says "That it's odd how we share results of RCTs (Randomised, Controlled Trials) in C19th essay format!" This is also how Cochrane Collaboration share reviews and meta-analyses of clinical trial data.


Structured data in RDF

Instead we should use "C21th structured data standards". I was especially pleased to hear how he was even more explicit: "Publish in RDF a good, quality standard, nice data format" [at 36.50 mins]

See also what the web development director at Cochrane, +Chris Mavergamessay in his excellent presentation on how linked data can help free content from the 'container of the article'.


This is related to our the work we do on linked clinical data standards, see my recent blog post: CDISC2RDF. That is, a semantic web versions of  data standards for clinical data on subject/participant level.

Clinical Data Transparency 

Given the recent move towards clinical data transparency (see a good summary in Nature this week Drug-company data vaults to be opened) I foresee a discussion also on data standards for the summary level data in clinical study reports and per-reviewed papers using semantic web standards.

An alternative could be to represent tables in the reports and paper as RDF using the RDF Data Cube Vocabulary (for multi-dimensional statistical data), see the CSVImport and the CubViz projects (Representing and browsing multi-dimensional statistical data as RDF using the RDF Data Cube Vocabulary, previously called Stats2RDF) This EU/FP7 project has used this vocabulary to publish biomedical statistical data, e.g. the WHO's Global Heath Observatory dataset (see Publishing and Interlinking the Global Health Observatory Dataset).

A challange is to express the clinical trial design and other contextual information as structured data to make it easier to make informed decisions for trial reviews and cross trial analyses.

Tuesday, February 12, 2013

CDISC2RDF

In a recent article from semanticweb.com (The Voice of Semantic Web Technology and Linked Data Business) the project CDISC2RDF is nicely decribed: Clinical Studies And The Road To Linked Data.

The project will be presented at the Conference on Semantics in Health Care & Life Sciences (CSHALS) meeting at the end of February by Charlie Mead, co-chair of the W3C’s Health Care and Life Sciences Interest Group (HCLSIG).

Here is a slide deck describing the first deliverable of the project. A refined slide deck will be presented at the CSHALS meeting together with a couple of CDISC2RDF blog post to describe the transformation process.

Saturday, December 29, 2012

My MOOCs Spring 2013

Great to see that the news program on SVT (Swedish Television) described MOOC (Massive Open Online Courses) in a new story the other day.

SVT Nyheter, 27 Dec. 2012:Toppuniversitet ger gratiskurser på nätet.  

During 2012 I have followed a few courses via one of the organisations mentioned in the news program: Coursera.Two of the courses were excellent: Model Thinking and  Fundamentals of Pharmacology, and they are on Coursera's list of 211 (!) courses. While the course in "Software Engineering for Software as a Service (SAAS") was not of the same high quality, and it's not on the list anymore.

For the Spring 2013 I have enrollod three MOOCs. So, now I know what to do while commuting 2 hours per day also the coming months :-)

It's great to see how all of this have taken off during 2012 offering courses not only for data nerds as myself but also for many others.

So, I was thinking of my sister when I read these teasers from Coursera:
  • "Ever wonder why people do what they do? This course offers some answers based on the latest research from Social Psychology."
  • "In the course Introductory Human Physiology students learn to recognize and to apply the basic concepts that govern integrated body function (as an intact organism) in the body's nine organ systems."

Sunday, September 16, 2012

Mind maps just begging for RDF triples and formal models

Earlier this week CDISC English Speaking User Group (ESUG) Committee arranged a webinar: "CDISC SHARE - How SHARE is developing as a project/standard” with Simon Bishop, Standards and Operations Director, GSK. I did find the comprehensive presentation from Simon, and his colleuage Diane Wold, very interesting.

Interesting as the presentation in an excellent way exemplifies how "Current standards (company standards, SDTM standards, other standards) do not current deliver the capability we require" Also, I do find the presentation interesting as it exemplifies mind maps as a way forward as "Diagrams help us understand clinical processes and how this translates into datasets and variables." (Quotes from slide 20 in the presentation: Conclusions.) 

Below a couple of examples of mind maps from the presentation. And also, the background to my thinking that they are Mind maps just begging for RDF triples and formal models of the clinical and biomedical reality to make them fully ready "both for human understanding and for computer interpretation".


High level mind map from the Parkinson's disease example
by Dianne Wold, GSK (slide 14)

Current standards do not current deliver the capability we require 

This conclusion is backed up in the first half of the presentation with exemples from GSK's internal standards and from CDISC's SDTM standards. These are low level data standard specifying data structures and data elements (variables). Standards for exchange of data in bulk (in containers such as SDTM Vital Sign and Lab domains) or standards for exchange of captured data (in specified variables such as a data modules for specific blod pressure and temperature mesurements) . Good exemples in the presentations show the challanges in analysing and aggregating clinical data put into SDTM dataset variables as containers "lacking documented relationships between the variables".

Example from data represented in the proposed
SDTM standard for Parkinson's disease (slide 12)

Diagrams help us understand clinical processes and how this translates into datasets and variables

The value in drawing diagrams to understand the higher level of relationsships in terms of the clinical processes in which clinical data is captured for different diseases. This is nicely illustrated in the presentation with a couple of diagrams, or "mind maps".

Example of a map of the clinical process (slide 15)
And also on the value in drawing diagrams to understand the mid level of relationsships in terms of "concepts" *) and "concept variables" and how these should be put into the SDTM variables (in red). (The example below is unfortunately not the same as the above Parkinson's disease examples.)

Example of a map for the concept of Temperature measurement (slide 28)


Mind maps just begging for RDF triples and formal models

When I see these mind maps I see graphs just begging for RDF triples (subject, predicat, object). That is, the fundemental semantic web standard. See my two earlier blog posts from two presentations at CDISC Interchange Europe: Semantic models for CDISC standards and metadata and Linking Clinical Data Standards

An intersting exercise would be to have the Parkinson's disease exemple completed in the concept mapping tool (CMAP) the whole way down to SDTM. And export the mind maps using as RDF triples. However, this is nice, but not enough ...

When I these mind maps I can also see how easy it is to start drawing such diagrams and exporting them as representations of generic mind maps. However, to fullfill the ultimate goal to have them "captured in a way that these can be used both for human understanding and for computer interpretation" the "mind maps" need underlying formal models of the clinical and biomedical reality.
 
Therefore, I see an interesting connection between the high level maps for disease and clinical processes to the Ontology for General Medical Science (OGMS). OGMS is an ontology of entities involved in a clinical encounter and provides a formal theory of disease that have been further elaborated by specific disease ontologies. See my blog post from last year on Disease terminologies and ontologies.


*) The CDISC SHARE project talks about scientific, or research, concepts. It has also been called observation concepts. However, the word "concept" is overused and carries challanges in itself, see From concept to clinical reality. 

Kudos to Frederik Malfait, working for Roche and my co-presenter on Semantic models for CDISC data standards and metadata, for pointing me to this presentation.