1.5
Thesis Contributions
This thesis contributes to the field of interacting with external displays in several ways. First, we extend the design space of ubiquitous mobile phone input with the aforementioned [BRSB08]. We then develop a new model for interacting with external screens by allowing users to inter- act with remote content through the live video image shown on their personal, mobile display. Iteratively, we extend this model to allow (1) non-modal connections, and (2) a fully flexible
environment. To meet the criteria of variable distances, we further extend the model to allow continuous operations with real-time input and feedback. Subsequently, we use the personal, mobile display not only as input surface but also use its visual output capabilities to give personal feedback. Additionally, we present an alternative that still allows all the presented interactions even though live video is not available. In the following we give a more detailed overview of the main contributions of this work:
Extension of the Design Space: Several approaches exist for interacting with external
displays using a mobile device. We give a detailed overview of the distinct phases within an interaction session as well as existing interaction techniques within each phase of a session. This analysis lead to an extension of the design space of ubiquitous mobile device input [BRSB08]. While the original taxonomy deals with the input domain, we extend it by means of the environment the mobile device is used in. The added factors are the flexibility of the environment, the variability of the distance between mobile device and target display, and the modality of the connection process. This design space is then used to create a new interaction model in fully flexible environments at varying distances without the need of explicitly connecting to a target display.
Non-Modal Connection Mechanisms: The increasing number of displays in a given
environment today requires users have to determine and select the display they want to interact with. When they interact with the display directly (e.g., by using touch in- put), this connection is established on a conceptual level in anon-modalway. However, performing such actions at a distance usually requires them to actively select a display on their mobile phone. Assuming autarkic displays in the environment, cross-display interactions can only be performed by establishing another connection to a second dis- play. By embedding the connection process into the interaction itself, we overcome this limitation from the user’s point of view. We propose a model that enables users to aim their mobile device (ands its built-in camera respectively) at the desired target display. They can then interactimpromptuwith the content shown in the viewfinder. At the same time, pointing at the display implicitly establishes a connection to the screen.
Flexible Environments: The more displays are present in an environment, the more
changes in the displays’ arrangement may occur. Especially when mobile displays are part of the environment (i.e., act as external display for a user), the environment is changing frequently and requires an interaction technique that allows for more flexi- bility. Thus, the non-modal connection procedure has to be enabled on all displays in the environment in the same way. We propose an infrastructure that realizes such
non-modal connection procedures on all displays. We further show how this can be used to allow for cross-display operations (such as dragging content from one display to another) without the need for disconnecting from one display and reconnecting to another. This is achieved by employing the mobile device’s camera (and its video re- spectively). Based on this video, the mobile device tracks itself with respect to displays in the environment. The presented infrastructure does not rely on pre-arranged displays or pre-modeled environments and allows forfull flexibilityof the environment.
Interacting at Variable Distances: When displays in an environment get larger, the
distance between them and users usually increases. Furthermore, multiple users may have different distances to the display. However, mobile devices usually feature wide- angle lenses which only allow for precise interaction at a maximum distance from the external screen. It is apparent that an interaction technique needs to support variable distances to allow for more flexibility of the user. We first allow users to zoom into the video image. With this, they can interact at higher distances with the same precision. The image further needs to be stabilized to counteract slight hand movements caused by, for example, the natural hand tremor. Despite modifying the existing infrastructure to allow real-time interaction for continuous operations, we address the mobility of the interaction device. We present three improvements to allow for a wide range of (meaningful) distances while maintaining operation accuracy and image stability.
Chapter
2
Related Work
We can’t solve problems by using the same kind
of thinking we used when we created them.
– Albert Einstein –The work described in this thesis aims to develop a novel model for interacting with external displays using a personal, mobile device. To inform the design space of such a model, this chap- ter gives a detailed overview of existing techniques. At the same time, we first focus on the three different phases of an interaction session - namely connectingto a display,pointingat its content, and manipulatingits content (see section 2.1). We identify the similarities and differ- ences compared to the ones found in traditional GUIs. According to these phases, we review existing research prototypes. We start with the logicalconnection that needs to be established between the mobile device and the external display before any interaction can take place (see section 2.2). Subsequently we review interaction techniques at-a-distance and pay close attention to both bringing content to the user as well as redirecting the user’s input (see section 2.3). To gain a deeper understanding of our design space, we then focus more detailed onpointingand in- teraction (i.e.,manipulation) techniques for external displays using mobile devices (section 2.4). At the same time, we investigate physical pointing (e.g., touching), relative pointing and pointing through the mobile device’s camera as interaction capabilities. In the following section, we re- view the interaction through video for both static (i.e., fixed mounted) and dynamic (i.e., mobile) cameras (see section 2.5). As our model aims to also allow for cross-display interactions, we also review existing approaches. At the same time, we focus on multi-display environments and describe existing techniques for exchanging content between the personal and an external screen or across two external displays (see section 2.6). We then summarize and discuss the presented approaches and arrange them in the extended taxonomy to identify unexplored regions (see sec- tion 2.7). The discussion and analysis then leads to our own exploration regarding different
2.1
The Phases of Interaction Sessions
Interacting with external displays usually requires users to establish a connection to the target screen: first, users need to identify the display they want to interact with (selection). Second, they need to establish a wireless connection to it. Subsequently, users can then control the ex- ternal display and manipulate content. These three steps are similar to traditional graphical user interfaces (GUIs) as described by Fitzmaurice: (1)acquire physical device, (2)acquire logical
device, and (3) manipulate logical device [Fit96]. In our context, the acquisition of a logical
device maps to the selection and connection to an external display using the personal, mobile display. The second step is to control a remote pointer which in turn allows the acquisition of content on the remote screen. The pointer control can be done using the rich sensing capabilities of today’s mobile devices. The subsequent manipulation can then be done by using the personal, mobile device (and its display in particular) as well.
Manipulate Select tool Acquire physical Device Acquire logical device Manipulate logical device Acquire logical Device Manipulate logical device Acquire external display Select tool Acquire content item Manipulate content item change tool same tool chan ge tool for sa me co ntent Interacting with external displays: Interacting in Traditional GUIs:
Figure 2.1: Interaction phases for both traditional GUIs and interacting with external dis-
plays: top shows the interaction phases for traditional GUIs. Users have to select the tool first and subsequently can position the pointer on the content in order to manipulate it. Bot- tom depicts the phases in our prototype. Users can select their tool (possibly eyes-free) while moving towards the content. This saves one phase and interaction time respectively.
However, before acquiring and manipulating content, users usually need to select a tool first (depending on the task). For example, in Photoshop, users need to select their brush before they can paint with the selected one in the image. Selecting the tool can also be expressed through these phases: acquire the tool and subsequently select (i.e., manipulate) it. Manipulating content required users to perform the second (i.e., acquire content) and third (i.e., manipulate content) twice for the overall goal of manipulating the content itself. The top of figure 2.1 shows the model for traditional GUIs with the extension of selecting a tool.
In contrast to traditional desktop computers, the input device in our prototype (i.e., the mobile device) features a display on its own. This display can be used to select the tool that is needed to perform the manipulation task. By doing so, users acquire the tool by selecting it on screen (i.e.,