Estimating Reliability of Telecommunication and Electronic Devices
Full text
Estimating Reliability of Telecommunication and Electronic Devices Compiled & Presented By: Engr. Dr. Muhammad Aamir
Reliability Reliability is defined as the probability that a device will perform its required function under stated conditions for a specific period of time 12/17/20152
Reliability Parameters… MTBF Mean Time Between Failures (MTBF), as the name suggests, is the average time between failure of hardware modules . between failure of hardware modules . It is the average time a manufacturer estimates before a failure occurs in a hardware module. 12/17/20153
Reliability Parameters… MTBF for hardware modules can be obtained from the vendor for off-theshelf hardware modules. MTBF for in house developed hardware MTBF for in house developed hardware modules is calculated by the hardware team developing the board. MTBF for software can be determined by simply multiplying the defect rate with KLOCs executed per second. 12/17/20154
Reliability Parameters… FITS (Failures in Time) FITS is a more intuitive way of representing MTBF. FITS is nothing but the total number of FITS is nothing but the total number of failures of the module in a billion hours (i.e. 1000,000,000 hours). 12/17/20155
Reliability Parameters… A correct understanding of MTBF is important. A power supply with an MTBF of 40,000 hours does not mean that the power supply should last for an average of 12/17/20156 supply should last for an average of 40,000 hours. An MTBF of 40,000 hours, or 1 year for 1 module, becomes 40,000/2 for two modules and 40,000/4 for four modules.
Reliability Parameters… MTTR Mean Time To Repair (MTTR), is the time taken to repair a failed hardware module. In an operational system, repair generally means replacing the hardware module. Thus hardware MTTR could be viewed as mean time to replace a failed hardware module. 12/17/20157
Estimating the Hardware MTTR … It should be a goal of system designers to allow for a high MTTR value and still achieve the system reliability goals. You can see from the table (Next Slide) that a low MTTR requirement means high operational cost for the system. 12/17/20158
Estimating the Hardware MTTR Where are hardware spares kept? How is site manned? Estimated MTTR Onsite 24 hours a day 30 minutes Onsite Operator is on call 24 hours a day 2 hours Onsite Regular working hours on week days as well 14 hours 12/17/20159 Onsite Regular working hours on week days as well as weekends and holidays 14 hours Onsite Regular working hours on week days only 3 days Offsite. Shipped by courier when fault condition is encountered. Operator paged by system when a fault is detected. 1 week Offsite. Maintained in an operator controlled warehouse System is remotely located. Operator needs to be flown in to replace the hardware. 2 week
Calculating Availability of Individual Components Component MTBF MTTR Availability Downtime Input Transducer 100,000 hours 2 hours 99.998% 10.51 minutes/ye ar Signal Processor 10,000 hours 2 hours 99.98% 1.75 hours/year 12/17/201516 Processor Hardware hours 2 hours 99.98% hours/year Signal Processor Software 2190 hours 5 minute 99.9962% 20 minutes/ year Output Transducer 100,000 hours 2 hours 99.998% 10.51 minutes/ year
Calculating System Availability Component Availability Downtime Signal Processing Complex (software + hardware) 99.9762% 2.08 hours/year Combined 12/17/201517 Combined availability of Signal Processing Complex 0 and 1 operating in parallel 99.99999% 3.15 seconds/year Complete System 99.9960% 21.08 minutes/year
Reliability Predictions Methods… There are generally two categories: (1) Predictions based on individual failure rates, and (2) Demonstrated reliability based on 12/17/201518 (2) Demonstrated reliability based on operation of equipment over time. Prediction methods are based on component data from a variety of sources like failure analysis, life test data, and device physics.
Reliability Predictions Methods… –MIL-HDBK-217 Generally associated with military systems –Models are very detailed –Provides for many environments –Provides multiple quality levels –Bellcore (Telcordia) Telecommunications Industry standard 12/17/201519 Telecommunications Industry standard –Models patterned after MIL-HDBK-217, but simplified –Provides multiple quality levels –Can incorporate current laboratory test data –Can incorporate current field performance data –Resources Software packages cover both MIL-HDBK-217 and Bellcore models –RelCalc MTBF Reliability Prediction Software (T-Cubed)
Reliability Predictions Methods For some calculations (e.g. military application) MIL-HDBK-217 is used, which is considered to be the standard reliability prediction method. 12/17/201520 calculations using Telcordia Technical Reference TR-332 “Reliability Prediction Procedure for Electronic Equipment.”
Typical MIL-217 Failure Rate Model A sample MIL-217 failure rate model for a simple semiconductor component is shown below. Many components, especially microcircuits, have significantly different and more complex models. Failure rate = pib * piT * piA * piR * piS * piC * piQ * piE 12/17/201521 Failures/million Hours Where: piT = Temperature factor piA = Application factor (linear, switching, etc) piR = Power rating factor piS = Electrical (voltage) Stress factor piC = Contact construction factor piQ = Quality factor piE = Operating environment factor
MTBF For Discrete LEDS and Photodiodes In the case of fiber optic data communication equipment, an area of primary concern is the MTBF values for the discrete Light Emitting Diodes (LEDs) and photodiodes used in the fiber optic transmitters and receivers . 12/17/201522 fiber optic transmitters and receivers . The determination of the predicted MTBF for an LED or photodiode is a calculation using the part failure-rate model for optoelectronic devices as found in section 5.1.3.10 of MILHDBK-217E
MTBF For Discrete LEDS and Photodiodes 12/17/201523
MTBF For Discrete LEDS and Photodiodes The Value for MTBF is then determined from: MTBF = 1/ λP Where: 12/17/201524 Where: TF= 8.01 x 1012 exp - (8111/TJ + 273) TJ= TA+ θJAPd TA is the Ambient Temperature Pdis the Power Dissipated by the Device
List of Constants For LED BF= Base Failure Rate = 6.5 X10-4 θJA = Thermal Resistance = 150 oC/W EF= Environmental Factor= 1 (Normal) QF= Quality Factor = 0.5 12/17/201525 For Photo Detector BF= Base Failure Rate = 1.1 X10-3 θJA = Thermal Resistance = 200 oC/W EF= Environmental Factor= 1 (Normal) QF= Quality Factor = 0.5
Ways to Improve Reliability Derating: Part failure rates generally decrease as applied stress levels decrease. Thus, derating, or operating the part at levels below it's ratings (for current, voltage, power dissipation, temperature, etc . ) can increase 12/17/201532 dissipation, temperature, etc . ) can increase reliability. Burn-In: Burn-in is operation in your factory, at elevated temperature, to accelerate the rate of failures; burn-in allows you to weed out failure prone devices in your factory, rather than in the field
Ways to Improve Reliability Redundancy: Product reliability may also be enhanced by using redundant design techniques. 12/17/201533
Thank You Thank You 12/17/201534