Showing posts with label evaluation. Show all posts
Showing posts with label evaluation. Show all posts

Monday, February 18, 2008

Multidimensional Evaluation Based on Story Telling and Scenarios

The interest in narratives has a long tradition. Bruner (1990) considers narrative a primitive function of human psychology, lying at the heart of human thought. The representation of experience in narratives provides a frame which enables humans to interpret their experiences. In this way narrative is a fundamental aspect in the construction of meaning.

Although narratives have been studied in many areas of psychology, the idea of exploiting them for system design and evaluation is quite recent and still not consolidated (Erickson, 1995). It is based on the need to understand user requirements through the collection of implicit knowledge that a user may have gained through experience. In particular, the use of narratives for system evaluation does not aim to provide quantitative results but to structure a framework within with users may express knowledge otherwise difficult to verbalize Stories make it possible to study the complexity of the context (media, internal/external environment, actors); details about critical events that can be observed very seldom; personal involvement; emotional details; and the concreteness and veracity of personal experiences.

Internet 2010

This chapter applies the narrative approach to the evaluation of an electronic tourist guide. The evaluation was carried out in an Italian museum, the Museo Civico, in Siena, with real visitors recruited at the museum entrance. The tourist guide is a prototype system developed within a project called Hyper Interaction within Physical Space (HIPS),' funded by the European Commission within the I-Cube (Is) Programme. The system is very advanced since it exploits cutting edge technologies (positioning technology, dynamic language generation) and visionary interaction design concepts (access to the information space through the physical movement in the museum; user modelling for content adaptation).

The system guides visitors by generating audio messages: users can get instructions on how to find items of interest, hear descriptions with references to items seen earlier and to ones that will follow, or ask for additional information. Information is generated dynamically, adaptive, and integrated with maps and spatial directions. When interacting with users, the system integrates their requests with a customized user-model, the user's browsing history, and their physical location at the moment of the query, providing highly contextual and personalized information (Marti et al, in press). The information content varies according to the user's location, preferences, and to the information already given. From a hardware point of view the system is based on a client-server model; the clients that visitors carry around are pen-driven palmtop computers with a screen, headphones and no keyboard. Localization is performed by various means: infrared, radio and GPS (Global Positioning System). Connectivity to the server is wireless.

Multiple Levels of Interaction, Multiple Levels of Performance Measure

The evaluation of such a system was quite complex because of installation problems within the museum's infrastructure, specific features of the system, and its particular context of use. For these reasons and in order to address the multi-faceted components of the HIPS system, we decided to adopt a multidimensional approach based on storytelling and scenarios. Specifically, we assessed user performance on four levels: phenomenological, cognitive, emotive and sociocultural.

At the phenomenological level the performance measure concerns included:

  • Users' perception of the adaptation to the visiting style (personalization of the information, pauses, pace of narration) and the physical movement as a primary means for accessing information.
  • Auditory comments (deictics, pronouns, etc) effectiveness in supporting the users' orientation and recognition of artworks.
  • Tool flexibility (skill of personalizing and contextualizing the information according to the users' changes of path or visiting style).
  • At the cognitive level the performance measure concerns focused on:
  • Internet 2010
  • The cognitive effort associated with the use of the tool and the comprehension of the contents.
  • Users'conceptual model.

At this level, scenarios were used to question the design of the system. We used Norman's cycle of cognition based on: goals, intentions, planning, execution, perception and evaluation to generate questions such as "How does the artefact evoke goals in the user?" or "How does the artefact make it easy or difficult to carry out the activity?" or "How does the artefact support the user when a shift in the goal occurs?".

At the emotive level the performance measure mainly concerned aspects of experiential cognition:

  • Observation of frustration.
  • Observation of confusion.
  • Expression of satisfaction.
  • Expression of engagement. At the socio-cultural level the performance measure concerns included:
  • The social aspects of group activity mediated by the system (communication, knowledge sharing, collective memories).
  • Appraisal /dislike of contents.
  • Impact of narrative styles (male/female voices, accents, music, reading styles).

Even if the narrative approach impacted the evaluation at the four above mentioned levels, it is worth clarifying that the stories mostly provided the structuring framework for reasoning on the system and externalizing implicit knowledge. For the evaluation of more detailed features of the system, from interaction design to contents, we complemented storytelling with other methodologies, including ethnographic observations of the activity and laboratory testing (heuristic evaluation, cognitive walkthrough, scenario-based evaluation on intermediate prototypes.

Saturday, February 16, 2008

Caught Between Real and Virtual Worlds

While agreement was breaking out at this workshop, so too were some concerns. As most HCI practitioners would agree the distinction between design and evaluation in HCI is often blurred, indeed design and evaluation have been described as the different sides of the same coin. The design of the DISCOVER system proved to be no different. As we elicited requirements and undertook early design we kept an eye on evaluation, ultimately as a sanity check. The use of the prototype, described above, had given us clear usability requirements. The affinity diagram had given us detailed requirements of the design, content and substance of the DISCOVER collaborative virtual environment but concerns as to the issue of evaluation (and with it validation) began to emerge.

Internet 2010

Design Tensions: No Magic, Thanks

For many collaborative tasks in virtual environments, the goal is to achieve a particular state of affairs within the virtual world, for example agreeing the layout of an office, as in Hindmarsh et al. (1998). For others, the activity within the CVE is part of a larger collaborative process, but the way collaboration works within the CVE need not exactly replicate real world interaction. In contrast, the DISCOVER environment must support the acquisition of real world skills. In short, skills acquired in the DISCOVER CVE must be transferable to the real world of the ship. This presents a significant constraint: ideally, interaction and collaboration must not be artificially harder or take longer than in the real world, but neither must they be artificially easier or executed faster. Thus the design of DISCOVER should treat with caution "magical" devices such as birds' eye views of the state of environment, Star Trek-like transporting, or visible rubber banding between an avatar and its current focus of attention.

Accordingly, much research which has addressed the problems of ensuring that the users of such environments are aware of their surroundings and of other users cannot be employed directly. Consider, for example, the problems with recognizing other avatars. To maintain realism, avatars cannot be simply labelled. There may also be the presence of dense smoke, and perhaps the need for the avatars to wear vision obscuring breathing apparatus. All of this means that recognizing one's colleagues (as avatars), mediated in the real situation by such characteristics as gait, stance and minor variations in standard issue clothing, becomes far more difficult.

Thus there is a fundamental tension between exploiting the technology to the full to produce a state-of-the-art virtual training environment and creating one which is faithful to the behaviour and constraints of a real ship. This, of course, is not merely a design issue but also presents a corresponding evaluation challenge.

Wednesday, February 13, 2008

Post-Test Session: Storytelling

At the end of each session, the visitors were involved in a debriefing (in one of the rooms of the museum) where storytelling was encouraged to further comment, analyse, and interpret events which occurred during the test. The subjects were asked to describe their experience looking at the video recording of the test. This created a common base of discussion and knowledge, and provided concrete data to express impressions and points of view. To facilitate the storytelling, we prepared a set of questions to stimulate the discussion at the four levels of evaluation. These questions spanned over the following dimensions:

  • Global experience
  • Satisfaction /engagement.
  • Contents.
  • Internet 2010
  • Design concepts.

As a general outcome of the evaluation, we can say that stories mostly addressed aspects related to the emotional and socio-cultural levels. Indeed the system created a rich sensory environment that the users perceived but were not able to describe in a more structured form than stories. Stories, being concrete and immediate, do not require abstraction or introspection, so the users were free to tell and to compare their previous experiences, re-creating a context to share with the facilitators. That is why we used storytelling as an expressive means to stimulate communication between people who were not familiar with each other and to encourage them to speak about personal feelings, experiences and impressions.

We were aware that this kind of evaluation cannot provide quantitative data about satisfaction and engagement (such as that obtained using the Differential Emotions Scale, the Semantic Differential Scale or the free labelling method (Kim and Moon, 1998)). However, since the system was oriented to entertainment and leisure, we decided to collect data that could give insights into the capability of the system to intrigue and attract the users. During the debriefing we collected a corpus of about 20 stories that were mapped on user requirements and system specifications. Here is one of those stories:

When we went to Avignone, to visit the Palais des Papes, we had a local guide, a teacher of a school party who explained in detail the artistic and historical features of every room. We were interested in her explanation, but the students (children of the primary school) got suddenly tired. During the visit we passed through a room, where there was a different kind of exhibition where strange and funny animals dangled from the ceiling. Pupils were very curious to know about them, but the teacher was prepared only on the Palais des Papes, so she passed through the room without paying any attention to the animals of the exhibition.

The story contains a number of relevant elements for the design and evaluation. It highlights that visitors have heterogeneous needs: most of the time their activity is "non-goal oriented" since they can be pushed just by curiosity or pleasure, and their behaviour is not predictable. They often do not know ahead of time, or with any specificity, what future state they desire to bring about. Therefore, the situations of use can be various and idiosyncratic, leading the visitors to frequently adjust their goals and objectives during the visit.

Taking these elements into account, we can infer that the HIPS tour guide and systems like it must learn to adapt to visitor inclinations as they arise. The solution designed into the system, which proved to be particularly appropriate in this respect, focused on establishing a closer link between environment and user interactivity through the introduction of adaptive mechanisms.

Design Tensions: Virtually Real or Really Real?

In addition to the technical challenges of designing and implementing this collaborative system, there are the more subtle issues of convincing senior professionals and their employers that the system is easy to use, is appropriately realistic and will deliver the training they require. Would you trust the captain of your passenger ferry who has practised his command and control skills on something which looks suspiciously like an arcade game? Similarly, would you, as personnel manager of a large shipping company buy this product and expect to have your ships' captains to use it?

The DISCOVER trainers believed that the collaborative virtual environment should aim to model a ship as fully and as accurately as possible to bridge this credibility gap. However, we recognize that the DISCOVER presentation of reality must necessarily fail, as it is likely that even small limitations will shatter the illusion. So the design challenge is to identify and abstract from reality those elements which will give a sufficiently good impression of a ship. But how real is real? To date, existing research has focused on achieving a sense of presence and evaluation instruments have been developed to measure just that, but what has not been established is whether "presence" is a good measure of being real. During our requirements work at one of the partners' sites in Denmark, a trainer told us that when a mariner was using their physical simulator they spoke of it as being "their ship" within 30 minutes of use. Physical simulators in contrast to a collaborative virtual environment are equipped with real, physical controls, readouts, charts and manuals with a synthetic display: for CVEs, all is synthetic.

Internet 2010

Evaluating Reality?

From a functional perspective, the environment must (for example) be robust, adequately fast, allow the movement of trainees and their interaction with each other and various objects, support trainer-trainee interaction, the modification of the environment by trainees and provide the numerous other functions specified in the requirements. Evaluation of such features is relatively simple, through inspection against the requirements list combined with simple trials covering the actions necessary to support the training scenarios. Narrow usability evaluation is again fairly unproblematic. Initially, we have used expert cognitive walkthroughs based on the structure suggested by the COVEN project, extended to cover aspects of usability for pedagogic interaction. These results are supplemented by the administration to representative users undertaking task-based trials of Kalawsky's VRUSE questionnaire instrument (Kalawsky, 1999), with additional material to elicit data about usability for collaboration. We are left with the questions, "However how does one evaluate reality?" and "Is presence an appropriate measure?". At the time of writing we are left with a round of iterative evaluation and redesign to determine whether the system is sufficiently real.

Internet Blogosphere