Sweet Sorrow Post-Mortem Interview
The following is a transcription of a conversation between:
- STEPHEN WITTEK (Shakespeare-VR project director, interviewer)
- JAEHEE CHO (Creative Director, Orta Interactive)
- BOKANG WANG (Programmer, ETC student)
- ANLAN YANG (Visual Designer, ETC student)
Links:
- Perception Neuron motion capture suit & software: (neuronmocap.com)
- Unity virtual reality game engine: https://unity.com/
- Autodesk Maya 3D computer graphics application: https://www.autodesk.com/
- Oculus Quest 2 virtual reality headset: oculus.com
SW. Okay, we’re recording! Bokang, please describe the process of making Sweet Sorrow, step-by-by-step, from your point of view.
BW. For the first month, I was mainly focusing on trying the new technology. I began by experimenting with the motion capture system.
SW. What is the name of that system?
BW. Perception Neuron. I also experimented with voice detection and sound detection software, but we didn’t end up using those technologies for the finished product. After that, we moved on to figuring out how to make the karaoke-style coordination of text and sound work in the game. By the second month, we had captured the motion-capture performances with the actors, so I started to process and clean up the animation. Then I imported the animation into the game, and tried to match the animation with the audio recording of the actors’ performances. [The actors’ physical and audio performances were recorded separately]. At the same time, I also had to make sure that the audio recording was in synch with the text-scrolling for the performers’ lines. Those three things—the animation, the audio, and the text—all had to work together at the same speed.
SW. Was that difficult?
BW. Yes, that was the most difficult part.
SW. So, just to back up a bit… We did the motion capture using the Perception Neuron suit and software—and that produced the animation. Then you exported the animation and imported it into Unity, correct? [Unity is a platform for creating virtual reality experiences, see unity.com].
BW. Yes. I imported the animation into Unity, and then Anlan combined it with the designs she created for the Romeo and Juliet characters.
SW. And all of the experimentation that you were doing with voice recognition and the text-scrolling… you were doing that work in Unity, correct?
BW. That’s right.
SW. So overall, what do you feel you have learned from your experience on the project?
BW. Well, first of all, I learned how to use the Perception Neuron motion capture system, which I hadn’t ever worked with before. Also, I got to know the balcony scene and the bedroom scene from Romeo & Juliet really, really well because I had to go over the lines hundreds and hundreds of times in order to get everything in synch. I have both scenes completely memorized.
SW. Ha! That was my secret objective all along: to turn ETC programmers into Shakespeareans… Okay, thank-you, very much Bokang. Anlan, you’re up next. Could you please describe the process that you went through as you designed the characters and the theater?
AY. Yes. So, first I created some conceptual sketches for the Romeo and Juliet characters.
SW. Were those hand-drawn sketches on pen and paper?
AY. I did them by hand, but not on paper. I used my iPad… After I had completed the concept art, I began to create the models using Autodesk Maya, a 3D computer graphics application. I started by focusing on color and texture for items such as Juliet’s dress. Then I worked through the process of rigging the characters.
SW. What does that mean?
AY. Well, we have to give the models skeletons so they can move like human beings. The rigging also enables me to refine features such as eye and lip movement.
SW. And you did all of that work in Maya, correct?
AY. That’s right. Then, once the models had skeletons, we were able to connect them to the animation data from the motion capture shoot. As Bokang mentioned, we had to do some clean-up work to smooth over the little glitches in the animation detail. We were aiming to make everything look as smooth and natural as possible.
SW. Let’s talk about your digital re-recreation of the Shakespearean theater. So, I know that you based your design on the 3D video that we shot at the American Shakespeare Company’s Blackfriars Theater in Stanton, Virginia. How did you go about adapting 3D video to create an interactive, 3D game space?
AY. Most of the theater modelling work—color, texture, layout, etc.—was also done in Maya. Then we imported the model into Unity, where we adjusted features such as shadows and the lighting from the candle chandeliers.
SW. Thanks, Anlan. I think that covers the basic details from the technical side of production. Jaehee, could you say something about the creative challenges?
JC. Well, when we started out, the first challenge was to figure out how achieve the basic goal of simulating a performance experience. There were a few different options. One option was to use speech recognition, as with something like Alexa or Siri. Another option was to use simple sound recognition, which recognizes noise but not words, as with a lamp you can turn on with a handclap. And yet another option was to follow something closer to a karaoke model, where the experience simply plays through irrespective of audio input. We were working on a fairly tight deadline, so choosing the most practical option was the first order of business. So as Bokang mentioned, we started out by doing some experiments with voice recognition and sound recognition. We wanted to build a really rough prototype, just so we could understand the basic mechanics of how the experience might work and get some feedback. Ultimately, we decided to go with the third option, which we have been calling the karaoke model. That was a really crucial creative decision, and a crucial turning point in the project.
SW. Bokang, what software were you using when you were experimenting with speech recognition and sound recognition?
BW. For sound recognition, we were using Unity, which has some functionality in that regard. For speech detection, however, we had to connect Unity to the AI website created by Meta. We designed Sweet Sorrow for the Oculus Quest headset, which is produced by Meta. They have a special server for AI speech recognition. So basically, the process would have involved recording the user’s speech—on a line-by-line basis—and uploading it to the server.
SW. Right. So, as Jaehee mentioned, we eventually decided to go with Option 3, the karaoke model, rather than using any sort of speech or sound recognition. What was the basis of that decision?
JC. Well, with speech recognition, you have to send a recording of the user’s speech to a server, so there is a little bit of a delay. The delay is not such a big deal if you’re simply issuing commands to something like Alexa, but it doesn’t work well for simulating a conversation, or a Shakespearean performance. For sound recognition, we don’t have to send anything to a server, so there isn’t a delay—but all it’s doing is recognizing sound. So with sound recognition, it doesn’t actually matter what the user says. Any sound—speech, background noise, or a car crash—could cue the next line… So, we had to weigh the pros and cons. At the end of the day, the karaoke model seemed like the best way to go.
SW. Any other challenges?
JC. Yes. As Bokang mentioned, another major challenge was the motion capture software, which was new to all of us. In retrospect, there are probably a few things I would have liked to have done differently. For example, in the balcony scene, the Juliet character moves around a lot, and it is difficult to maintain eye contact. If I could go back, I would ask the performer to remain in a more-or-less stationary position during the motion capture shoot, so it would be easier to orient the character toward the user’s perspective in virtual reality.
SW. Are there any other lessons that you want to keep in mind for future projects?
JC. Well, aside from the technical lessons, I was really reminded of the importance of subtle, creative details. In the final weeks of the project, we started to adjust aspects such as lighting, sound design, music, and little creative details. At that point in the production, the experience really started to come alive. Very minor details, such as the sound of the user’s footsteps, can make a big difference to imaginative engagement.
SW. I know exactly what you are talking about. In those final few weeks, the experience suddenly started to work. We went from, “okay, this works” to “hey, this is fun!”
JC. Exactly. I think the audience made a difference to the overall feel of the experience as well. We originally tried to create a 3D audience, so the theater would be populated by characters along the same lines as Romeo and Juliet. But somehow, it just didn’t work. It was just a little bit off, and seemed like too much of a distraction. So ultimately, we ended up using a more abstract 2D impression of an audience adapted from the 3D video that we shot at the Blackfriars Theater in Staunton. It’s very simple, but I think it’s just enough to give the user a sense of an audience without being intrusive.
SW. Yes, I recognized the audience. I remember meeting that guy wearing a tie in the middle front row…
JC. In most cases, virtual reality experiences don’t use two-dimensional artwork—which makes sense, of course, because the whole point of virtual reality is to provide an immersive, three-dimensional experience. But in this case, it worked really well.
SW. The moment when Romeo and Juliet kiss during the bedroom scene is a similar example, right? We knew that Romeo and Juliet had to kiss, but we weren’t sure how to do it.
JC. Yes. We couldn’t make the 3D models kiss because they can’t get that close to each other without distorting the graphics. The solution we came up with was to blend 2D artwork within the 3D experience in order to create an abstract, impressionistic representation of the moment that would still work on an emotional level. It was a bit of a risk, but I think it worked.
SW. Yes, I think it’s a really fun, cute moment. Paradoxically, despite all our technological sophistication, we ended up using a lot of the same techniques that were very common on the Shakespearean stage, where a very simple prop could create the impression of something more fully realized in the spectator’s imagination.
