All systems run in degraded modes, and this means failures can be difficult to predict. Therefore, adaptability is key for dealing with unexpected events. Adaptability supports resilience. Both reliability and resilience are needed for ever increasing speed and scale. Read the full article at ieee.org
Read moreThe Resilience Engineering Association (REA) celebrates the 20th anniversary of the creation of Resilience Engineering, as a field of theory and practice, by strategically reflecting on progress made, challenges and future opportunities. To this end, the Resilience Engineering Association will publish a book that will collect contributions on progress and challenges and opportunities in the […]
Read moreThe introduction of Artifical Intelligence into software operations can have unintended consequences on software safety & reliability. Closer examination and realistic expectations of human and machine capabilities, along with thoughtful interaction design can produce safe, productive human-machine teams. Read the full article at ieee.org
Read moreIn the last Failure Mode article, we looked at how observability and explainability are critical—but insufficient—design elements of a reliable human-machine team. In this article, we connect ideas and research from cognitive systems engineering into high-impact design principles for improving safety and reliability of highly automated and intelligent systems. Read the full article at ieee.org
Read moreFor resilience in complex and large-scale software systems, we need to go beyond observability and explainability and consider joint cognition between human–machine teams. Read the full article at dl.acm.org.
Read more