Charlotte Ward: 0:14
Hello and welcome to episode 255 of the Customer Support Leaders podcast. I'm Charlotte Ward Today. Welcome Lauren Rose Eimers to talk about handling outages.
Charlotte Ward: 0:33
I'd like to welcome to the podcast today. Welcome back, actually, after quite a long time. What a hiatus for both of us in some ways, I guess. But please welcome back, Lauren Rose Eimers. Lauren, it's lovely to see you again. We have caught up a couple of times this week and it's been an absolute delight to talk to you after, I think, we figured out more than a year. But welcome back. For the benefit of everyone else, would you like to introduce yourself and hi?
Lauren Rose Eimers: 0:59
Well, hello, and thank you so much for having me back. I've missed being on the podcast. I've missed the podcast, so I'm thrilled to be back and thrilled to also learn from other leaders that have been sharing their time recently. But again, my name is Lauren Rose Eimers and I have 10 plus years now in the tech sector as a customer success and customer support individual. I've worn all the hats working at a startup from higher number I don't know, I think seven or eight, I can't even remember watching the company grow and supporting that through many different roles and ultimately, I now consider myself a customer support leader. So I'm happy to say after a decade of that, but I'm still learning. I think that's the most important thing is that you never stop learning, especially in this industry with all of the changes and everything going on. So I'm thrilled to be here today and chat a little bit with you about how we mitigate outages or unstable releases or unreliable products when you're on customer support or success.
Charlotte Ward: 2:04
Oh my goodness. First of all, yes, I'm still learning too, and I know that I've been in this game significantly longer than you have. Despite that decade of experience, I'm still learning too. So, yeah, I think that's why we're all here having these conversations. There is so much. Every time I have a conversation with somebody on the podcast, I come away with something new, and one of my very good friends, Craig Stoss I'll call him out because I'm sure he's listening out there somewhere I like to think he is Craig says you never leave a conversation with less information, which is a catchphrase of his that I took to heart very early on in our friendship and I've carried that with me for the last several years. I love that view on all of these conversations, so it's definitely a learning experience every time we talk. Welcome back. So outage management, handling outages, handling buggy software, bad releases, all of that fallout. Where do you want to begin with this conversation?
Lauren Rose Eimers: 3:09
Well, first I want everyone to knock on wood, because we don't even utter that word without fearing that we're going to bring that upon ourselves, because it's never a fun time. I think the first thing is to name outages. Of course, by their very nature unplanned and nobody likes unplanned things. When you're logging in on a Tuesday morning to thousands of tweets or emails or what have you alerting you that there's instability in your product, so there's all sorts of different ways that you could go about mitigating this. But I kind of like to look at this from a framework that could be used for anything from an acute or a very short outage, which many software products we like to just call them hiccups. Right, like it's a hiccup, thank goodness, everything's back online.
Lauren Rose Eimers: 3:56
Two longer standing outages, things like DDoS attacks, or a build was released that just for some reason has a huge bug that is not going to be fixed until the next build is released, to truly long standing, where it's baked into the product, if you will, and I wouldn't call that necessarily an outage, but basically a feature that needs mitigation until an upgrade can be completed, right.
Lauren Rose Eimers: 4:24
So I think the very first step is identifying the issue. So if you cannot recreate what your customers are reaching out to you about, that's a problem because, with any sort of outage, you want to be able to be communicating with your product team in real time, instead of giving anecdotal evidence or playing that game of telephone where somebody whispers something in your ear and then you're whispering that to the devs. No, you and your team need to be figuring that out, kicking the tires yourself, to be able to recreate the problem. This is also helpful because, especially for those really acute outages where it's like system-wide I mean thousands, millions, even of customers are being affected you want to be able to distill information down into need to know which, when people are angry about a product that's unstable, they're not giving you need to know information. They're giving you a lot of emotion, which is completely normal, and they want answers.
Charlotte Ward: 5:22
They're not coming to you. A lot of technical noise as well. I often try and draw parallels with, like, maybe supporting granddad trying to accomplish a task on his iPhone or something. Oh, my goodness, I don't know about you, but I'm technical support for the family when granddad's on the phone trying to get something done on his phone and I'm at cross-town and I'm just like, just take it slowly, try one thing at a time and talk me through exactly what you're doing, but he's jabbing at other bits of the screen and telling me about irrelevant information. It signals to him noise, to me as a technical support person quite often. So this message, that message, that color, that button, and actually I'm just interested in trying to, as you say, distill the problem to its really core essence. Get that accurate description so you can reproduce it and get that need to know information back to product over to the next guys.
Lauren Rose Eimers: 6:25
Yes, I love that image of distillation. You're given a lot of this raw material, but as a support individual, you are the filter in which that distillation needs to occur. You're recreating the issue, distilling information and reporting that, hopefully in real time, to the devs on the product team that are now doing step two, which is mitigating the problem, before any outage updates, any mass tweets go out or messages to your users, I think it's imperative that you are able to say that something is being worked on and also given a loose timeline, because that's the main thing folks want to know, especially if something is really truly not working or it's a huge blocker to their own use of the tool. I mean, I think we forget sometimes. People are using our software because it's solving a problem for them.
Lauren Rose Eimers: 7:21
When you stop solving that problem, that's an issue and you need to be able to say we are working on this problem and this is the expected amount of time for this problem to be solved. I think that's imperative. Of course, that can change right, so that's why, I mean, I love status pages for this reason, because that's when you can update, let folks know this is in progress, this is being worked on, this has been mitigated. Those are all things that are incredibly important, which brings me to my number three step, when it comes to dealing with outages, which is communicate.
Charlotte Ward: 7:59
Oh, my goodness, so important. In every direction, right, this begins with the problem identification through to mitigation. But as a support team, you're in the middle. You've got to be able to communicate in both directions effectively. Part of that mitigation cycle that you were just talking about, it is understanding what promises you can make to the customers coming out of our product or engineering teams. How do you do that?
Lauren Rose Eimers: 8:39
Well, that's why I said communicate three times, because you need to be communicating with the devs, your product team, and that has to be happening, again, in real time. You need to be communicating with your team, allocating a certain person on your team to answer tickets in the queue, another person to handle social media, another person to update PagerDuty. Communicate with your team because many times an outage, this is the first rodeo for some of your newer members. This is something new to them. So communicating to them and you helming that during a pretty stressful outage is imperative. And then, of course, the last communicate is with your customers and again, I mentioned this before, you know having status page updates. If you have in-app messaging banners that can help with those kind of things, I shy away from sending out mass emails because sometimes outages only affect a certain pocket of users, correct?
Lauren Rose Eimers: 9:35
Very true. I think you know, taking to social media. Those are the kind of things where you're communicating to your users and then, of course, having an outage macro. If you're using a tool that you can send off individualized messages to folks that are writing in, with again the fact that you have identified the issue, it's being worked on, you're working as hard as you can and this is the expected time things will be back online. That, again, is something that needs to be updated and again that might change. So the three communicates: you need to be communicating with product, your own team and externally with your customers, so that in any times of crisis, overcommunication is key.
Lauren Rose Eimers: 10:17
You want everyone to know what's going on, and so really take to the tools that your team uses. If it's in Slack, maybe you have a dedicated outage Slack channel. If everyone wants to hop on a Google Meet call or a Zoom call and co-work in real time if you're remote, leaders can be helpful as well. If you're back in the office, having forbid, you are all maybe head to the same conference room so you can work in real time with each other. Those kinds of comms are so helpful in those acute stressful times.
Charlotte Ward: 10:50
And I think an important part of the comms if I can just chip in on that a little bit as well is that it helps you build a single source of truth, doesn't it? It helps everyone be on the same page, ensuring that nobody's going off making up their own kind of worldview and saying to a subset of customers this is what we think, these are the promises we're going to make, and you've not communicated that to the wider team internally. That's where those macros and that consistent story internally really really are crucial. Actually, because you can't be sending mixed messages to different groups of customers and, even worse, to the same customer, but from two different sources. Maybe a status page there's one thing and a support agent says something else. How infuriating as a customer. So I think building that single worldview internally is really important.
Lauren Rose Eimers: 11:47
Oh, I couldn't agree more, and I think this is where, as a side note, standard operating procedures, your SOPs, really come in handy. Having an outage protocol that folks can harken back to, make sure they're not being forced to make choices in a really heated or stressful time that they shouldn't have to make right. I think good management of a team, good leadership of a team, allows there to be frameworks that folks can pull from in times of stress.
Lauren Rose Eimers: 12:14
You don't have to have all of the answers in the moment, but if you have a framework to operate from, it's so much easier to build off of that and so letting your teammates know from day one hey, we have these SOPs that you can pull from if you ever have any questions. Lean on your teammates as well, but please don't be sending out these rogue emails, especially during times when we're all kind of stressed out that might be going against the current, if you will, when you're trying to send out a unified singular message.
Charlotte Ward: 12:47
Yeah, absolutely.
Lauren Rose Eimers: 12:49
We're gonna say, all right, everything's been mitigated, wonderful. Yay, this was just a hiccup, it's not longstanding. But the next step of this framework is gratitude. I know that sounds really weird and a little hippie-dippy, but I will say thank your customers, your users, for hanging in there with you, even if it was a 15 minute hiccup. You send an email to each and every user that reached out to you. You send reply tweets to every single person that took to social media, every DM, and you thank them for their patience and thank them for letting us know, letting you know that something was wrong.
Lauren Rose Eimers: 13:28
So the gratitude there, but also gratitude for the experience, and every single outage is an opportunity to learn and refine. So you need to be grateful for that experience. I know it sounds a little Pollyanna, but you truly can take so much from every situation. Just like we talked about at the opening, if you weren't learning from an outage, which is basically a conversation with your product that was unstable for a certain amount of time, that's a problem. So the next step, after that gratitude, is learn from the situation.
Charlotte Ward: 14:03
I'm sorry to keep pulling you back to the previous step, though.
Lauren Rose Eimers: 14:07
Oh no go ahead.
Charlotte Ward: 14:09
We're definitely moving on, I promise, but I want to sit in gratitude for a minute because I just think that's such an important thing and we don't have that really on any incident plan that I've ever seen. The thing that I would add to that is please sound human while you're doing it. Oh yes, nobody wants to hear we apologize for the inconvenience. Nobody wants to hear that.
Lauren Rose Eimers: 14:37
Right, right. Oh, I think gratitude is such an opportunity to create a really personal connection. Yes, and the way I kind of think about the customer experience, be it in support or success. Every time a customer reaches out to you is an opportunity to delight them, to put a little money in that bank of, I trust you, I like you, you're people and I like using your product, and I think that that's not quantifiable, it's something that I harken back to the phrase people won't necessarily remember what you say, but they will remember how you made them feel and if you can make folks feel heard, seen and also attended to in a way that's human, it's going to do amazing things for your customer support and success team and as well as your customer base.
Charlotte Ward: 15:32
I'm gonna add a little more into that as well, which is that there is I don't know if you've heard of the peak end rule, but there is this idea in any interaction, and certainly this is true for customer experience which is that the two things your customers will remember about any experience with you are the peak of their emotions.
Charlotte Ward: 15:53
So they will remember how happy or how angry or how engaged and delighted they were when those emotions were at their height, whether that's at the beginning or the middle or the end. But they'll remember how that experience finished as well. So they will, for instance, remember how angry you made them two minutes into this incident, but they will remember how you finish it as well. And that actually is super important in the rescue, because if they can balance those two points the peak and the end, then you can overcome a lot of early anger in any incident with how you finish it, and that connection is crucial to making that end positive and memorable. And this moment where you're thanking customers is, I'm going to guess, the moment where most of the customer interaction in an outage finishes.
Lauren Rose Eimers: 16:45
It does, unless I think there again, every situation is different, but for these hiccups, personally, I think that's where the interaction can end and it be appropriate. If this is a long-term issue, if this is something that customers have been writing in for months about, I think updates are imperative. When the thing is actually fixed, when the new build is finally released, you save the folks that have truly reached out and have been impacted by these issues. That's when you close, when you say, hey, it might have taken us X amount of time, but this is now completed and we want to make sure you've upgraded or updated to the new build. I think that there are again more opportunities to repair. But I've never heard of that phrase. I'll have to look into it more. But it makes so much sense what I know about how brains work and the way neurons fire along with emotion.
Charlotte Ward: 17:46
It makes so much sense with that peak, intense emotion, but also the ability at the end. I wrote about it, or at least I wrote an article that drew on some inspiration from it. I'll send you the link to that, and I will include it in the show notes for this episode too.
Lauren Rose Eimers: 18:01
Thank you. See, I'm learning so much already from you. Both of us, all the time.
Lauren Rose Eimers: 18:06
Yes, this has been a generative conversation. Oh, and I forget, learning from the experience was where I left off, but that's pretty self-explanatory. If you're not having follow-ups to go over what happened, how it occurred, how folks communicated, and also those growth edges I like to call them, kind of arching back from my counselor days, but those growth edges where maybe we fumbled a bit here and there. But this is where we can shore up our SOP, this is where we can do some more preemptive training to let folks know how to deal with this in-house.
Lauren Rose Eimers: 18:41
There's so many ways that things can be honed and improved upon. Again, nobody wants an outage. Everything we do is in prevention of outages but in the event that it does occur, just like I have a background in the medical world but you practice things over and over and over again in hopes you never see them in clinic. But having that framework to operate from can help you during those higher times of stress and it's very similar to outages. Finally, just to close everything up, is to adjust the way you deal with it, moving forward.
Lauren Rose Eimers: 19:19
That adjustment is imperative. You can learn and you say I learned so much. But if you don't make adjustments on your end, when the next outage happens to come around hopefully not for many moons, but these are things that again, you're honing, you're learning, but then you're adjusting for the next time so that either it's an adjustment on the product end, where maybe they do more QA before releasing a build, or maybe they were able to find the area in the code base that needed shoring up, or there were inconsistencies or vulnerabilities in certain areas of the product. All of these things are the need for adjustment across the board. So all of that to say, outages are no fun, and so if you've gone through one recently, I'm so sorry.
Lauren Rose Eimers: 20:07
I also think this is not just the framework that you need to operate from, but also thank your team. Let them have a little like send them a free cup of coffee or their favorite beverage. Or, if it's a long running outage and they've had to deal with it for a few days, there's nothing wrong with sending them a gift certificate to a local spa, or give them a day off or two, understanding that it can take a toll and that stress, especially in the moment and having to be on and handling high stress situations can really take a toll on your team. So gratitude for your team is imperative. But that's at the end of all of this, when hopefully, everything's been mitigated, we've learned from it, we've adjusted and now your teammates get to step away from the keyboard and have a little time to rest and recoup, because I think sometimes it's lost on us just how much emotional labor customer support and success folks engage in daily and then to add things like the stress of an outage, it's just, it's a lot.
Charlotte Ward: 21:14
It's a lot, it really is.
Lauren Rose Eimers: 21:16
Yeah, I think allowing your team a little space and grace to recoup is also imperative, because you need to take care of the folks that are taking care of your customers.
Charlotte Ward: 21:25
You need to decompress, don't you. Right? Yeah, I mean just that decompression is necessary for all of us, but the focus on the frontline looking after these situations, and for leaders too, who are managing all of the operational side. But you know, there's a lot of politics around this. In my experience as well, and not necessarily in the hiccups, but in those big ones, yeah, I mean it's worth just jumping on. If the least you can do is jump on a Zoom and just say, phew, well done, guys, that goes a long way.
Lauren Rose Eimers: 22:05
It really, really does. Recognizing the extra effort and the extra load that that places on you. Emotionally it's a lot. So and yeah, thanking the team and not everyone has to be perfect in times of crisis. I think we have to remember that perfection is never the goal. I mean it would be wonderful if we could be perfect in those times of stress, but I think transparency, humanity and empathy can work wonders in those times of stress and, honestly, that kind of counts for all areas of our life. It doesn't have to necessarily be in software, but I mean.
Charlotte Ward: 22:42
For me, that is a key aim of the last two points of your framework, that retrospective and then follow up actions, which is something we do. For every incident, every outage, we run the, we try and run the true blameless retrospective, so nobody is at fault. And the thing I'm always saying is, we don't fix people, we fix the systems. You know so the people aren't to blame. If things go wrong, if they couldn't do something or if they did something wrong, you know wrong in inverted quotes, if they were able to put us on a path that made this situation worse in any way, the system is at fault. And so we fix the system for next time so that we don't, so that we always get, we're always heading towards the happy path and not away from it. You know that's the ideal right.
Lauren Rose Eimers: 23:37
Well, I love that lens that you're looking through. I think that your team is lucky to have you at the helm, but also I think that, again back to, I love how this all just came full circle, but learning from the experiences and having your team learn alongside, where no one is at fault, but also making those adjustments for down the road, when those acute, stressful situations may arise.
Charlotte Ward: 24:00
Absolutely, absolutely. Thank you so much, Lauren. So much to learn there, as always, and so much food for thought. Thank you so much for joining me after such a long break. It was great to have you back and nice to catch up with you this week. Please come back and have another conversation soon, won't you?
Lauren Rose Eimers: 24:17
I absolutely will. This was a pleasure and again, I hope me saying outage so many times doesn't bring that upon any of us. This is all hypothetical. No one's going to have an outage ever again. It's going to be fine.
Charlotte Ward: 24:30
All fine, all good. Thank you so much.
Lauren Rose Eimers: 24:32
Thank you so much, Charlotte. It was a pleasure.
Charlotte Ward: 24:37
That's it for today. Go to customersupportleaders.com forward slash 255 for the show notes and I'll see you next time.