{"status":"ok","message-type":"work","message-version":"1.0.0","message":{"indexed":{"date-parts":[[2026,6,10]],"date-time":"2026-06-10T16:00:51Z","timestamp":1781107251122,"version":"3.54.1"},"reference-count":13,"publisher":"IGI Global Scientific Publishing","issue":"3","content-domain":{"domain":[],"crossmark-restriction":false},"short-container-title":[],"published-print":{"date-parts":[[2011,7,1]]},"abstract":"<p>Failure detection is a fundamental building block for ensuring fault tolerance in large scale distributed systems. It is also a difficult problem. Resources under heavy loads can be mistaken as being failed. The failure of a network link can be detected by the lack of a response, but this also occurs when a computational resource fails. Although progress has been made, no existing approach provides a system that covers all essential aspects related to a distributed environment. This paper presents a failure detection system based on adaptive, decentralized failure detectors. The system is developed as an independent substrate, working asynchronously and independent of the application flow. It uses a hierarchical protocol, creating a clustering mechanism that ensures a dynamic configuration and traffic optimization. It also uses a gossip strategy for failure detection at local levels to minimize detection time and remove wrong suspicions. Results show that the system scales with the number of monitored resources, while still considering the QoS requirements of both applications and resources.<\/p>","DOI":"10.4018\/jdst.2011070105","type":"journal-article","created":{"date-parts":[[2011,10,20]],"date-time":"2011-10-20T10:38:27Z","timestamp":1319107107000},"page":"64-87","source":"Crossref","is-referenced-by-count":10,"title":["A Failure Detection System for Large Scale Distributed Systems"],"prefix":"10.4018","volume":"2","author":[{"given":"Andrei","family":"Lavinia","sequence":"first","affiliation":[{"name":"University Politehnica of Bucharest, Romania"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Ciprian","family":"Dobre","sequence":"additional","affiliation":[{"name":"University Politehnica of Bucharest, Romania"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Florin","family":"Pop","sequence":"additional","affiliation":[{"name":"University Politehnica of Bucharest, Romania"}],"role":[{"vocabulary":"crossref","role":"author"}]},{"given":"Valentin","family":"Cristea","sequence":"additional","affiliation":[{"name":"University Politehnica of Bucharest, Romania"}],"role":[{"vocabulary":"crossref","role":"author"}]}],"member":"2432","reference":[{"key":"jdst.2011070105-0","doi-asserted-by":"crossref","unstructured":"Bertier, M., Marin, O., & Sens, P. (2002). Implementation and performance evaluation of an adaptable failure detector. In Proceedings of the International Conference on Dependable Systems and Networks, Edinburgh, Scotland (pp. 354-363).","DOI":"10.1109\/DSN.2002.1028920"},{"key":"jdst.2011070105-1","doi-asserted-by":"publisher","DOI":"10.1145\/226643.226647"},{"key":"jdst.2011070105-2","doi-asserted-by":"publisher","DOI":"10.1109\/12.980014"},{"key":"jdst.2011070105-3","doi-asserted-by":"crossref","unstructured":"Costan, A., Dobre, C., Pop, F., Leordeanu, C., & Cristea, V. (2010). A fault tolerance approach for distributed systems using monitoring based replication. In Proceedings of the 6th IEEE International Conference on Intelligent Computer Communication and Processing, Cluj-Napoca, Romania (pp. 451-458).","DOI":"10.1109\/ICCP.2010.5606398"},{"key":"jdst.2011070105-4","article-title":"A dependability layer for large scale distributed systems.","author":"V.Cristea","journal-title":"International Journal of Grid and Utility Computing"},{"key":"jdst.2011070105-5","unstructured":"Defago, X., Hayashibara, N., & Katayama, T. (2003). On the design of a failure detection service for large-scale distributed systems. In Proceedings of the International Symposium on Towards Peta-bit Ultra Networks, Ishikawa, Japan (pp. 88-95)."},{"key":"jdst.2011070105-6","unstructured":"Dobre, C., Pop, F., Costan, A., Andreica, M. I., & Cristea, V. (2009). Robust failure detection architecture for large scale distributed systems. In Proceedings of the 17th International Conference on Control Systems and Computer Science, Bucharest, Romania."},{"key":"jdst.2011070105-7","unstructured":"Dobre, C., Voicu, R., Muraru, A., & Legrand, I. C. (2007). A distributed agent based system to control and coordinate large scale data transfers. In Proceedings of the International Conference on Control Systems and Computer Science, Bucharest, Romania."},{"key":"jdst.2011070105-8","doi-asserted-by":"crossref","unstructured":"Fetzer, C., Raynal, M., & Tronel, F. (2001). An adaptive failure detection protocol. In Proceedings of the 8th IEEE Pacific Rim Symposium on Dependable Computing, Seoul, Korea (pp. 146-153).","DOI":"10.1109\/PRDC.2001.992691"},{"key":"jdst.2011070105-9","unstructured":"Hayashibara, N., Defago, X., & Katayama, T. (2004). Flexible failure detection with K-FD. (Tech. Rep. No. IS-RR-2004-006). Ishikawa, Japan: Japan Advanced Institute of Science and Technology."},{"key":"jdst.2011070105-10","doi-asserted-by":"crossref","unstructured":"Nastase, M., Dobre, C., Pop, F., & Cristea, V. (2009). Fault tolerance using a front-end service for large scale distributed systems. In Proceedings of 11th International Symposium on Symbolic and Numeric Algorithms for Scientific Computing, Timisoara, Romania (pp. 229-236).","DOI":"10.1109\/SYNASC.2009.13"},{"key":"jdst.2011070105-11","doi-asserted-by":"crossref","unstructured":"Stelling, P., Foster, I., Kesselman, C., Lee, C., & von Laszewski, G. (1998). A fault detection service for wide area distributed computations. In Proceedings of the 7th IEEE Symposium on High Performance Distributed Computing (pp. 268-278).","DOI":"10.1109\/HPDC.1998.709981"},{"key":"jdst.2011070105-12","doi-asserted-by":"crossref","unstructured":"Van Renesse, R., Minsky, Y., & Hayden, M. (1998). A gossip-style failure detection service. In Proceedings of the Middleware Conference, The Lake District, UK (pp. 55-70).","DOI":"10.1007\/978-1-4471-1283-9_4"}],"container-title":["International Journal of Distributed Systems and Technologies"],"original-title":[],"language":"ng","link":[{"URL":"https:\/\/www.igi-global.com\/viewtitle.aspx?TitleId=55422","content-type":"unspecified","content-version":"vor","intended-application":"similarity-checking"}],"deposited":{"date-parts":[[2022,6,2]],"date-time":"2022-06-02T00:24:56Z","timestamp":1654129496000},"score":1,"resource":{"primary":{"URL":"https:\/\/services.igi-global.com\/resolvedoi\/resolve.aspx?doi=10.4018\/jdst.2011070105"}},"subtitle":[""],"short-title":[],"issued":{"date-parts":[[2011,7,1]]},"references-count":13,"journal-issue":{"issue":"3","published-print":{"date-parts":[[2011,7]]}},"URL":"https:\/\/doi.org\/10.4018\/jdst.2011070105","relation":{},"ISSN":["1947-3532","1947-3540"],"issn-type":[{"value":"1947-3532","type":"print"},{"value":"1947-3540","type":"electronic"}],"subject":[],"published":{"date-parts":[[2011,7,1]]}}}