Showing posts with label data model. Show all posts
Showing posts with label data model. Show all posts

Monday, March 5, 2012

Is a Definition Just a List of Attributes?

If we look at a data model, is a definition of an entity type automatically produced by listing the attributes of the entity type?  If this were true then a data modeler would not need to produce entity definitions - he or she would simply need to identify and list a sufficient number of attributes.  I have actually heard data modelers being criticized by terminologists for doing just this.  The extent to which such criticism is fair or not is a separate discussion, but the question remains as to whether a list of attributes can suffice as a definition.

I do not think that a list of attributes is sufficient based on the recent discussions about concept systems in this blog.  No concept exists in isolation.  Every concept exists in some kind of concept system where it has relationships to other concepts.  At least some of these relationships and/or related concepts have to enter into a definition so that the concept being defined can be located properly in a concept system, which appears to be necessary for knowledge.

A list of attributes usually will not distinguish between those that are determining for the concept under consideration, and those which are not - some of which may be shared with other concepts.  Thus, just reviewing a list of attributes becomes a test of figuring out which ones are pertinent to a definition.  This surely defeats the practical aspects of definition.

It would seem therefore that something more than a list of attributes is required to produce a quality definition.  While many data modelers do produce quality definitions, it can be seen that the practice of data modeling may present the temptation to just assume that the definition of an entity type is provided by the attributes captured for it.  Of course, relationships and other concepts are present in a data model, but it will need another blog to answer the question of whether a data model has enough information to produce a definition based on entity types, attributes, and relationships alone.

Monday, February 6, 2012

The Idea of Concept Systems

An involuntary hiatus has prevented me from the pleasure of blogging on definitions for about a month.  I am now gradually getting back to normal, and am able to blog again.

Today I want to look at concept systems, and types of concept system.

In data modeling, only one type of concept system commonly appears - the generic concept system, containing Supertypes and Subtypes.  Very occasionally, the part-whole type of concept system can also be found.  The latter be seen in "bill of material" structures.  Strangely, the visual representation of a generic concept system and a part-whole concept system can look very similar in a data model.   I think that this leads data modelers to play down the idea of concept systems, and indeed the term "concept system" is not really met with in data modeling.

However, if we turn to the discipline of terminology, the idea of concept system is very prominent, and different types of concept systems are called out.  Let me quote from the Nordterm Guide to Terminology by Heidi Suonuuti (ISBN 952-9794-14-2):

"Concepts are not independent phenomena.  They are always related to other concepts in one way or another, and form concept systems which can vary from fairly simple to extremely complicated.  In terminology work, an analysis of the relations among concepts and an arrangement of them into concept systems, is a prerequisite for the successful drafting of definitions."

It is interesting that from the data modeler's perspective, concept systems are viewed only with respect to designing data storage solutions.  A terminologist, by contrast, is more interested in business information and how concepts are related within it - irrespective of how such information might be stored as data.

This makes me wonder about semantic modelers.  We hear a lot about semantics these days, and there is no doubt that semantics involves identifying concepts and providing definitions for them.  But finding the relationships between the concepts must be done prior to forming the definitions.  This is because a definition, in part, describes a concept's relations to other concepts in the concept system in which it is found.  So what good methodologies, notations, and techniques exist for describing or visualizing concept systems?  I am not sure we have yet got any good ones.   The danger is that we then fall back on the data modeling methodologies, notations, and techniques, which fail to capture significant semantic details.

But perhaps more important is that the terminologists have the idea of types of concept systems.  The generic and partitive types of concept system are the major ones, but there are others.  We will deal with the different types in a future post.

Tuesday, December 27, 2011

Is A Data Model An Abstraction?

Rob brings up a good point in his comment on The Problem of Abstraction in Definitions of Data (http://definitionsinsemantics.blogspot.com/2011/12/problem-of-abstraction-in-definitions.html).  He notes that what I am describing is not really abstraction but really a number of different things.
Today it seems the term "abstraction" is used in all kinds of situations when talking about data.  For me, it is often difficult to figure out what "abstraction" is supposed to mean in any one of these situations.  I strongly suspect that at least sometimes it does not really mean anything.  Sometimes I suspect it is even used for marketing hype.

The entry for "abstraction" in Baldwin's Dictionary of Philosophy and Psychology describes how abstraction is filtering out of attributes from an instance or a concept to achieve a particular view of the instance or concept.  Rather poetically the entry describes how a child looks at a body of water and becomes fascinated by the lustre caused by the play of sunlight on the surface of the water, to the exclusion of all the other qualities (attibutes) of the water.

This traditional understanding of abstraction as creating a view by filtering out attributes can be used in a special way to create the generalization hierarchies of genus and species (a.k.a. supertype and subtype, or general concept and specific concept).  The particular attributes of a group of specific concepts are left behind and attributes that the concepts have in common remain.  These are used to form the general concepts that include the specific concepts. 

However, abstraction as filtering out of attibutes can generate other perspectives.  Abstraction does not always have to lead to the traditional generalization hierarchy.  I can understand a man's watch as a timepiece, or a piece of jewelery, or as a fashion accessory.

Now, Rob is right in that I was not using "abstraction" in the above senses.  However, I do not have a better term to use for what I was trying to describe.  The main idea I was trying to get across is that one concept system can describe or specify another - such as how a  data model describes a physical database.  The relationship of "description" here is different to every other kind of relationship because the concepts present in the concept system being described have to have some kind of presence in the concept system doing the the describing.  This is not the same class of relationship we see in e.g. "I own a car".

So we somehow have the presence of a concept being described (e.g. a column of a physical database table) in a concept system doing the describing (e.g. an attribute of an entity type in a data model).
Rob terms this "representation" (If I understand his comment correctly).  This has to be right.  However, a representation can often be a picture - a mere image.  Technically, this is called a "phantasm" because it does not have the attributes differentiated from the whole.  Unfortunately, the process of recognizing and separating the attributes from a phantasm is also called abstraction.  It gets more complicated.  we cannot take a photograph of a physical database and produce anything like a data model.  A database has to be conceieved, not imagined.

Obviously, we are getting into a whole lot of other issues here.  I cannot really defend myself against Rob's criticism of my having overloaded (or over-abstracted?) the term "abstraction".  However, I do not have a commonly accepted set of terms that I can use to convey the idea of one concept system describing another.  More of an excuse than a reason, but it will have to do for now.

Wednesday, December 21, 2011

The Problem of Abstraction in Definitions of Data Objects

I think there is a major problem in not being able to understand and work with different levels of abstraction.  By "abstraction" in this sense I mean one concept system that somehow describes or defines (not merely relates to) another concept system.  I think this is a big problem for definitions in data models.

Let us take an example in a retail business such as mortgage banking: Customer Name.  Customer Name exists in the business.  They use it all the time.  Maybe it is sometimes called Borrower Name, but the concept is the same.  This is the Level 1 abstraction.

Now let us think of data values in a column in a table that holds Customer Name.  These data values are stored as a code of 1's and 0's.  Of course these bits are rendered into something we can read.  However, this is not the same as the Customer Name in the business.  I worked for a place where they prefixed the name of anyone who had recently left with "ZZZ".  So we could have "ZZZ_John Smith" as a data value, but the business would call him "John Smith" still.  The data value is the Level 2 abstraction.

Now let us think of the column itself that stores Customer Name, irrespective of whatever it contains.  This is the container used for the data.  It is merely a container, and anything can be put into it - just in the same way as the old peanut jelly jar I have on my desk is used to hold pens.  The column has certain characteristics, like the maximum length of text it can hold.  This is the Level 3 abstraction.

Now let us think of the data model that describes the column that will hold Customer Name.  In this, Customer Name is an attribute.  We worry about what naming convention to give it.  And behold!  Our data modeling tool asks us to enter a definition for Customer Name!  Yet, we are now at Level 4 of abstraction.

Let's summarize.  The concept system of the data model (Level 4) is a design for the concept system of the container of the data (Level 3) which will store the concept system of data values (Level 2) which we hope will satisfy the concept system of the information needs of our users (Level 1).

So tell me again what the definition entered in the data model is referring to?  Which of the four concept systems?.  Suppose it is stated as "an attribute that holds customer name" - I have seen this kind of thing quite often.  Well, an attribute is something in a data model (Level 4), and a thing that holds data is a container (Level 3).  

It would seem that the ideal thing would be to understand the Level 1 abstraction - the business information.  However, the chances of getting a good definition of this when you are at Level 4 would seem to be a challenge.  There are too many layers of abstraction in the way.  This, I think, is why semantics are so important.  They deal with business information as is, and do not have to worry about other concept systems.