Saturday, November 26, 2011

Paper Reading #5: A Framework for Robust and Flexible Handling of Inputs with Uncertainty

References
Julia Schwarz, Scott Hudson, Jennifer Mankoff, and Andrew D. Wilson.  "A framework for robust and flexible handling of inputs with uncertainty". UIST '10 Proceedings of the 23rd annual ACM symposium on User interface software and technology.  ACM New York, NY, USA ©2010.

 Author Bios
Julia Schwarz is a Ph.D. student at Carnegie Mellon University.

Scott Hudson is a professor at Carnegie Mellon University, where he is the founding director of the HCII PhD program.

Jennifer Mankoff is a professor at the Human Computer Interaction Institute at Carnegie Mellon University.  Mankoff earned her PhD at the Georgia Institute of Technology.

Andrew Wilson is a senior researcher at Microsoft Research.  Wilson received his Ph.D. at the MIT Media Laboratory and researches new gesture-related input techniques.
Summary 
  • Hypothesis - The researchers hypothesized that using the uncertainty aspect of a gesture during the process to interpret it will give more accurate results.
  • Method/Content - They created several examples to test their framework. It works by sending the user input to all possible recipients of the gesture, along with the information about it. When it sends this along, it also calculates the probability of it being the correct choice. However, the evaluation system used in this framework is lazy; that is, it waits until the last possible second to actually decide which action to take. Once the gesture or gesture sequence is completed (noted by specific actions, such as taking a finger off the screen or a specific length period of silence), it then by probability decides which action to take. One of the examples they used was voice recognition where 'q' and '2' sound the same.
  • Results - All of their tests came out positive. The findings were that their system did increase the accuracy of interpretation by a large margin. The results from four movement impaired subjects with conventional motion gesture recognition was that between 5 and 20% of all inputs were interpreted incorrectly. However, with the new system, less than 3% for all four of them were misinterpreted.
  • Content - The researchers first discuss the limitations of current gesture search systems. They put forth their claim that getting rid of uncertainties in the beginning of the system greatly increases the chances of misinterpreting a gesture. The then created a system in which it keeps these uncertainties and the associated information. Using this new system, they test several participants and compare the new system with the current one. They hypothesis was supported, and they go on to discuss future possible applications.
 Discussion
The thoughts of the developers were that rather than getting rid of uncertainties in the beginning when they have nothing else to go on, they should wait to get rid of them until the end when something inputted later may narrow down the choices. This is a surprisingly new concept that has not been strongly implemented yet. Especially in today's world where screens are getting smaller and touch-interaction is the norm, being able to select actions based upon probability is a must. This would be a great thing to integrate into many systems in use today, like smart phones and tablet PCs.

Paper Reading #4: Gestalt

References
Kayur Patel, Naomi Bancroft, Steven M. Drucker, James Fogarty, Andrew J. Ko, and James Landay.  "Gestalt: integrated support for implementation and analysis in machine learning". UIST '10 Proceedings of the 23rd annual ACM symposium on User interface software and technology.  ACM New York, NY, USA ©2010.

 Author Bios
Kayur Patel is a PhD student at the University of Washington specializing in machine learning.

Naomi Bancroft is a Senior undergraduate researcher at the University of Washington. Her interests lie in HCI.

Steven M. Drucker is a  Principal Researcher at Microsoft Research with who specializes in HCI. He received his PhD from MIT.

James Fogarty is currently an assistant professor at UW. His research focuses on HCI and Ubiquitous computing. He received his PhD from Carnegie Melon.

Andrew J.Ko is also currently an assistant professor at UW. His research focuses on the “Human aspects of software development”. He also received his PhD from Carnegie Melon.

James A Landay is a professor at UW, whose research focuses on Automated Usability Evaluation, Demonstrational Interfaces, and Ubiquitous Computing. He received his PhD from Carnegie Melon as well.

Summary 
  • Hypothesis - A general purpose Machine Learning tool that allows developers to analyze the information pipeline will lead to greater efficiency and fewer errors.
  • Method/Content - The researchers created two problems, one for movie reviews and one for gesture recognition. Eight testers were then given a program for each problem; each program had 5 bugs in it. Within an hour, they were asked to find and fix as many bugs as they could. The tools they used to find and fix these problems were their newly developed Gestalt Framework, and the other was to use a customized version of Matlab. Each participant was asked to solve each problem with each program (4 tests in all).
  • Results - The results showed that participants were able to find significantly more errors while using the Gestalt framework. Some even tried to create Gestalt functionality within Matlab. All eight of the users preferred Gestalt over Matlab, and most of them stated that they would likely benefit from using Gestalt in their work. 
  • Content - This paper presented Gestalt, which is a new tool for developers of Machine Learning. It then conducted a user study to compare it with other similar software, and found that it was indeed a good tool. It then discussed its strengths and weaknesses. Its main strength lies in its ability for users to view the information pipeline.
 Discussion
Although a general purpose tool cannot necessarily perform all of the same tasks as well as a domain-specific tool, they are often flexible enough to still be a powerful tool. Gestalt seems as if it has a good ways to go before it sees general use, but the results were promising. Throughout the paper, the greatest thing that I saw about the framework was its ability to let you view (and manipulate) the information pipeline. This is key for many applications, especially machine learning. Although their testing methods were not robust, they did serve to show a general sentiment of how Gestalt can be useful to developers.

Wednesday, October 19, 2011

Paper Reading #23: User-Defined Motion Gestures for Mobile Interaction

References
Jaime Ruiz, Yang Li, Edward Lank.  "User-Defined Motion Gestures for Mobile Interaction". UIST '11 Proceedings of the 2011 annual conference on Human factors in computing systems.  ACM New York, NY, USA ©2011.

 Author Bios
Jaime Ruiz is a fifth-year doctoral student in the HCI Lab at the University of Waterloo. His main research interests is further understanding users to augment the level of human computer interaction. For this paper he was a research intern at Google Research.

Yang Li is a senior research scientist at Google Research. Before joining Google, he was a research associate in the Computer Science and Engineering at the University of Washington, where he helped found the DUB. He received his Ph.D. in Computer Science from the Chinese Academy of Sciences, and conducted postdoctoral research in EECS at the University of California at Berkeley.

Edward Lank is an assistant professor at the University of Waterloo. Before he joined Waterloo, he was an assistant professor at San Francisco State University and conducted postdoctoral research at Xerox. He received his Ph.D. in Computer Science from Queen's University.

Summary 

  • Hypothesis - Conducting an end-user design test will result in best practices for designing gestural interfaces.
  • Method/Content - This paper focused on a user study involving motion gestures using mobile phones. Twenty participants were involved. Because learning about a new device would have cause people to be less receptive to new concepts, they required that the participants used a smartphone as their primary mobile device. Nineteen tasks were given to the participants. For each task, the participant was asked to come up with an easy and comfortable gesture to conduct the desired action. For example, to come up with a gesture for answering a call. Because they did not not want users to focus needlessly on recognizer issues or the current level of technology, they asked participants to treat the device as a "magic brick", capable of automatically understanding and recognizing anything they might throw at it (figuratively, of course). The participants were asked to conduct the study while thinking aloud. In other words, they wanted the users to explain what they were doing, what they may be emulating, and why for each gesture. Once the participant had decided on a unique gesture for each of the 19 tasks, they were asked to perform each one five times. While they conducted each gesture the phone recorded data from the accelerometer and its other motion-detecting hardware and sent to another computer. Once they were done, the users were asked to rate their own gestures on a Likert scale based on its ease of use and whether or not they would use it often.
  • Results -The results were that the closer a gesture was to emulating a real world object, the more of a consensus among participants was reached. For example, when answering a call, easily the most common gesture was to simply raise their phone to your ear. This makes sense, as for a normal call, this would be the first thing you do (usually after you answer it). An example of emulation that does not happen regularly is the act of hanging up the phone. The gesture users found the most consensus on this one was to emulate an "old-fashioned" phone; that is, they turned the phone so that its screen was parallel to the ground. Another thought that many users had is that some of the actions should have the same gesture; for example, going to the next photo, contact, or search results, it should all be a flick to the right. Users generally agreed that although in different contexts, the end result was basically the same, so the same gesture should be used. For actions that did the opposite of another, users generally performed the same gesture, only in the opposite direction. For example, going to the previous item in a list required a flick to the left. Another example is zooming in and out on a map; zooming in was to bring it closer to your face, zooming pushing it farther away.

    Another curious result is that, in contrast to surface gestures, users generally wanted to move the window as opposed to the object. This means that when dealing with a map or image, the users panned left by moving the phone to the left, whereas with a surface gesture it is moving to the right. This was explained by the fact that when touching the screen itself, the user was in effect moving the object on the screen. However, when moving the phone, the user was moving the screen around the object, and expected it to react as such.
 Discussion
This paper seemed like it is a bit late in coming, although I enjoyed it. It seems that we have had the technology for a long time, and it took us this long to start looking for motion gestures? I remember back when the Gameboy Color was popular; Kirby Tilt 'n' Tumble was one of my favorite games. That used an accelerometer, which is one of the main tools used in motion gestures. It just surprised me that it took us this long to make the transition. However, I think that this technology/concept has a large range of practical applications. Things like conducting presentations or interacting with colleagues in a design room would be great places to use a device for this. However, this could also raise the problem of using a phone in the car, which we already have enough issues with. As seen in Why We Make Mistakes, making the phone accessible without looking at it would not solve the problem. Humans simply cannot multitask that well.

Thursday, September 29, 2011

Paper Reading #13: LightSpace

References
Andrew D. Wilson, Hrvoje Benko.  "Combining Multiple Depth Cameras and Projectors for Interactions On, Above, and Between Surfaces". UIST '10 Proceedings of the 23rd annual ACM symposium on User interface software and technology.  ACM New York, NY, USA ©2010.

 Author Bios
Andrew D. Wilson is a senior researcher at Microsoft Research. He received his bachelor's from Cornell, and subsequently his master's and Ph. D. at MIT. He helped found the Surface Computing group at Microsoft.

Hrvoje Benko is also researcher at Microsoft. He received his Ph.D. from Columbia University. His  interests revolve mostly around augmented reality and in discovering new ways to blur the line between 2D computing and our 3D world.

Summary 
  • Hypothesis - This paper did not have a hypothesis; it was simply a discussion and description of a design concept.
  • Method/Content - The main concept behind this design was to allow users to interact with projected displays using normal tables (or other ordinary flat surfaces) as touch-screens. It made use of a suspended apparatus containing 3 IR and depth cameras and 3 projectors. The device decided what the user was doing by creating a 3D mesh of them and simulating their movements in a virtual 3D space. It created the mesh by using its cameras to create depth maps from different angles. Because it used the notion of one virtual space for all 3 cameras and projectors, there was no real discrepancies or major mistakes in gesture recognition. Available interactions included dragging some projected object off of a table, holding said object, putting the object back on the table (or a different one), moving objects from table to vertical screen and back, transferring objects from one person to another, and an interactive menu. The menu worked by moving through options based upon the height of the users hand, and selecting it if the user held it there for 2 seconds. Holding an object worked loosely on the idea of holding a ball; the user held their hand level and carried the "ball" where the wanted it to go. They could let go and drop the ball at any time.
  • User feedback on this system was overall very good. During their public demonstration they discovered a number of limitations that were not readily apparent, but were possible to fix. One of the limitations was that if there were more than 6 people in the room, the system would get confused as to who was who because everyone was too close to one another. It was hard for the system to distinguish what person was trying to perform what action. This is easily enough fixed by increasing the space in the room and the range of the cameras/projectors. Another limitation found was that having 3 or more people in the room slowed the system drastically; the refresh rate of the system (and thus the projectors) dropped below the camera's refresh rate. This is also pretty easily fixed, to a point: use a more powerful computer to render the 3D space and meshes used. There is obviously an upper bound on the amount of people the system can accommodate (due to both size constraints and computing power) but it can definitely perform better than what has already been implemented.
 Discussion
Again, I loved this paper. It seems that the more I read, the more I realize that we are much closer to virtual reality environments than I thought. This system has a huge range of applications, from meetings to showcases, from artists to engineers, from product design to video games. The concept of moving objects from one surface to another is not really what excites me; it's the system itself. The fact that they can use relatively simple and inexpensive cameras to track multiple entities without the users wearing external apparatus (ie dots or markers) is amazing. I would absolutely love to have this system in my house, if just to play around with and maybe customize (to perform different actions).

Tuesday, September 27, 2011

On Gangs

Sudhir's Gang Leader for a Day was a very interesting book. I thoroughly enjoyed all but the last few pages. In those pages, I was really disappointed that he said they had never been friends. It seems to me that when you go through that many things together for that long and still enjoy being around one another, you have become friends. Even to this day, Sudhir visits JT whenever he is in Chicago; that says friend to me. I suppose he needed to say that (according to lawyers) in order to acquit himself and show that he is and was not associated with any gangs, but it still seems to be a terrible way to end the book.

During the book, it astounded me to see what the people living in the projects did and endured to survive, especially the women. Some of the things they did I had previously associated with third world countries; that it was happening here in the US and especially in one of the most influential cities made me sad. Although thinking about it now, I suppose it should not have come as a surprise; you'll find poverty anywhere.

When Sudhir was a gang leader for a day, I felt that he embellished on a lot of it. I think that his decision about the guy who stole and the guy who withheld pay was correct; but it still wasn't his decision in the end. Throughout the day he was sort of riding shotgun instead of driving; JT would do most things and occasionally ask what Sudhir thought, then took it as advice rather than as instruction. There are reasons he couldn't truly make any of the final decisions: JT couldn't afford to lose face in front of his subordinates, if he made a wrong decision it could cause the loss of a lot of money, etc. But I still don't think 'gang leader for a day' is a proper description for what he did; 'gang leader adviser for a day' is a much more apt description.

When the projects were torn down, I felt bad for JT and his two long time friends. Yes, they were gang leaders and yes, they were perpetuating the use of drugs, but they truly believed that what they were doing helped the community as a whole (or so they claimed). Although they were perhaps rough about it and obtained the money for it through unethical means, they did what they needed to to survive; with those methods they also helped the community in many ways, whether or not they had an ulterior motive for it. I felt really sad when T-bone died; he truly had a plan for after the gang life. He wanted to get a degree, live normally and honestly. He seemed to me one of those that truly got caught up in something they didn't want and couldn't get out.

All in all, I really enjoyed the book. Sometimes sad, sometimes happy, but most of the time just interesting. The end was disappointing, but that by no means made it a bad book. I would definitely recommend this book for the future classes. 

Sunday, September 25, 2011

Paper Reading #12: Enabling beyond-surface interactions

References
Thomas Augsten, et al.  "Enabling Beyond-Surface Interactions for Interactive Surface wit An Invisible Projection". UIST '10 Proceedings of the 23rd annual ACM symposium on User interface software and technology.  ACM New York, NY, USA ©2010.

Author Bios
Li-Wei Chan is a Ph. D. student in the Graduate Institute of Networking and Multimedia at the National Taiwan University. He received his master's and bachelor's in Computer Science from the National Taiwan University and from Fu Jen Catholic University respectively.

Hsiang-Tao Wu, Hui-Shan Kao, and Home-Ru Lin are students at the National Taiwan University.

Ju-Chun Ko is a Ph. D. student at the Computer & Information Networking Center, National Taiwan University. He got his master's in Informatics from Yun Ze University.
Mike Y. Chen is a professor in the Department of Computer Science at National Taiwan University. His research interests lie in mobile technologies, HCI, social networks, and cloud computing.

Jane Hsu is a professor of Computer Science and Information Engineering at National Taiwan University. Her research interests include intelligent multi-agent systems, data mining, service oriented computing and web technology.

Yi-Ping Hung is a professor in the Graduate Institute of Networking and Multimedia at National Taiwan University. He received his bachelor's from National Taiwan University and his Master's and Ph.D. from Brown University.


Summary
  • Hypothesis - Using IR (infrared) cameras to place invisible markers will improve reliability for interactive tabletops.
  • Method -For this experiment, they used a custom interactive tabletop prototype. It projected both color and IR from under the table, and used two IR cameras under the table to detect touches. The IR projector also selectively projects white space on the tabletop to perform multi-touch detection. The tabletop itself is comprised of two layers: a diffuser layer and a touch-glass layer. Due to the reflective nature of the touch-glass, it caused problems whether it was above or below the diffuser layer. They found that when it was above, it reflected the visible light of projections from above the tabletop, which caused not only a degrade in the luminance of the projection, but also shined the light on observers. When the glass was under the diffuser layer, it partially reflected the IR rays from beneath the table, resulting in dead zones for the image processing. They found that they could fix the dead zone problem by using two IR cameras instead of one, so they implemented the table with the touch-glass underneath the diffuser layer. The IR cameras used a dynamic sizing system to track projections and move/resize markers as needed. The proposed 3 different projection systems: the i-m-Lamp, the i-m-Flashlight, and the i-m-View. The first was a combination pico-projector/IR camera which appeared as a simple table lamp. Its small dimensions were thought to be ideal for integration with personal tabletop systems. The second (i-m-Lamp) implementation proposed is a mobile version of the i-m-Lamp. Users can inspect fine details of a region by focusing the i-m-Flashlight at the desired location. The i-m-View is a tablet PC attached to an IR camera. The programmed use for it was to intuitively explore 3D geographical information. They used the i-m-View to explore 3D buildings from above a 2D map shown on the prototype tabletop system. They asked 5 users to try out their systems and were encouraged to think aloud.
  • The main problems found for the i-m-Lamp was that because the i-m-Lamp and the tabletop system both project on the same surface, the overlapped region caused a blue artifact. To avoid it, they masked the tabletop projection where the projections overlapped. For the i-m-Flashlight, they encountered a focus problem; the lens focus of the pico-projectors needed to be manually focused. This limited usability; however, they proposed that replacing the projector with one that contains a laser (such as the Microvision ShowWX) would provide an image that is always in focus. The largest problem with the i-m-View was that it was easy for users to get lost in the 3D view and not be able to pay as much attention to the 2D map. They fixed this by showing the boundaries of the 2D map inside the 3D view, allowing the user to simultaneously see what was changing on the table and what it represented in the 3D view. During for the i-m-View users often found that the buildings in the 3D view were too tall for the view; they wished to either pan up or rotate the tablet in order to get a portrait view of the landscape, neither of which were currently supported by the system. Another problem was that they i-m-View occasionally got lost because no IR markers entered its field of view; this was dealt with by continuously updating the orientation of the i-m-View. The overall feedback from users was positive, and the problems discussed are supposed to be addressed in future work.

    Discussion
    Quite frankly, I found this entire paper awesome. I thought that much of it was quite advanced, a huge step in HCI. While it may not have much application for me personally (I can't readily see this augmenting programming in many ways), it would have huge impacts on artists, the military, modelers and designers, engineers (such as civil or mechanical) and many more. Artists could use it to selectively edit only certain portions of their work without using the cumbersome selection methods used in today's art development programs. The military could quite easily use this for strategic purposes such as battle maps or location coordination. Modelers and engineers could use this to select certain pieces in a 3D model or blueprint to edit. In short, this technology has a huge range of applications that would make great use of it. I hope to see this technology distributed widely soon.

    Saturday, September 24, 2011

    Paper Reading #11: Multitoe

    References
    Thomas Augsten, et al.  "Multitoe: high-precision interaction with back-projected floors based on high-resolution multi-touch input". UIST '10 Proceedings of the 23rd annual ACM symposium on User interface software and technology.  ACM New York, NY, USA ©2010.

    Author Bios
    Thomas Augsten, Konstantin Kaefer, are a master student of IT systems at Hasso Plattner Institute (University of Potsdamn) in Germany.

    Christian Holz is a Ph. D. student in Human Computer Interaction at the Hasso Plattner Institute. He believes the only way to continue to further miniaturize mobile devices is to fully understand the limitations of human computer interaction. 

    Patrick Baudisch is a professor in Computer Science at the Hasso Plattner Institute.

    Rene Meusel, Caroline Fetzer, Dorian Kanitz, Thomas Stoff, and Torsten Becker are students at the Hasso Plattner Institute.


    Summary
    • Hypothesis - Using foot input is an effective way to interact with a back-projected floor based computer.
    • Method - The first study conducted was intended to be built off of for subsequent experiments. It was to test how buttons could be intentionally walked over without activating them. Participants were asked to walk over 4 buttons, two of which were meant to be activated, 2 of which were not. User methods were recorded and categorized. The second study determined which area of the foot user expected to be detected to activate a button. The third study was to determine if there was consistency in preferred hotspots across the user base. The fourth was meant to determine user ability; they were asked to type a sentence using a projected keyboard.
    • The results for the first test were that users did not generally have a consistent way to activate buttons. In the second test, most users agreed that the foot's arch was the best way to activate a button. The third test showed that users had virtually no agreement between users; no hotspots had the majority of usage. In the fourth test, it was found (as expected) that the smaller the keyboard, the more errors the user made. Users were about even in their preferences of the medium and large keyboards.
    Discussion
    While I'm not sure that this technology has immediate application, I believe that this could be one of the first steps to virtual reality rooms. I really enjoyed the concept, although I'm not sure that the users enjoyed it as much as me. Current uses may be exploring maps, or games such as Dance Dance Revolution, or if they include multi touch (with a large amount of possible touches) it could support group activities or games.