934 resultados para Test reliability
Resumo:
OBJECTIVE: The aim of the study was to validate a French adaptation of the 5th version of the Addiction Severity Index (ASI) instrument in a Swiss sample of illicit drug users. PARTICIPANTS AND SETTING: The participants in the study were 54 French-speaking dependent patients, most of them with opiates as the drug of first choice. Procedure: Analyses of internal consistency (convergent and discriminant validity) and reliability, including measures of test-retest and inter-observer correlations, were conducted. RESULTS: Besides good applicability of the test, the results on composite scores (CSs) indicate comparable results to those obtained in a sample of American opiate-dependent patients. Across the seven dimensions of the ASI, Cronbach's alpha ranged from 0.42 to 0.76, test-retest correlations coefficients ranged from 0.48 to 0.98, while for CSs, inter-observer correlations ranged from 0.76 to 0.99. CONCLUSIONS: Despite several limitations, the French version of the ASI presents acceptable criteria of applicability, validity and reliability in a sample of drug-dependent patients.
Resumo:
BACKGROUND: Anterior shoulder stabilization surgery with the arthroscopic Bankart procedure can have a high recurrence rate in certain patients. Identifying these patients to modify outcomes has become a focal point of research. PURPOSE: The Instability Shoulder Index Score (ISIS) was developed to predict the success of arthroscopic Bankart repair. Scores range from 0 to 10, with higher scores predicting a higher risk of recurrence after stabilization. The interobserver reliability of the score is not known. STUDY DESIGN: Cohort study (diagnosis); Level of evidence, 2. METHODS: This is a prospective multicenter (North America and Europe) study of patients suffering from shoulder instability and waiting for stabilization surgery. Five pairs of independent evaluators were asked to score patient instability severity with the ISIS. Patients also completed functional scores (Western Ontario Shoulder Instability Index [WOSI], Disabilities of the Arm, Shoulder and Hand-short version [QuickDASH], and Walch-Duplay test). Data on age, sex, number of dislocations, and type of surgery were collected. The test-retest method and intraclass correlation coefficient (ICC: >0.75 = good, >0.85 = very good, and >0.9 = excellent) were used for analysis. RESULTS: A total of 114 patients with anterior shoulder instability were included, of whom 89 (78%) were men. The mean age was 28 years. The ISIS was very reliable, with an ICC of 0.933. The mean number of dislocations per patient was higher in patients who had an ISIS of ≥6 (25 vs 14; P = .05). Patients who underwent more complex arthroscopic procedures such as Hill-Sachs remplissage or open Latarjet had higher preoperative ISIS outcomes, with a mean score of 4.8 versus 3.4, respectively (P = .002). There was no correlation between the ISIS and the quality-of-life questionnaires, with Pearson correlations all >0.05 (WOSI = 0.39; QuickDASH = 0.97; Walch-Duplay = 0.08). CONCLUSION: Our results show that the ISIS is reliable when used in a multicenter study with anterior traumatic instability populations. There was no correlation between the ISIS and the quality-of-life questionnaires, but surgical decisions reflected its increased use.
Resumo:
BACKGROUND This study assesses the validity and reliability of the Spanish version of DN4 questionnaire as a tool for differential diagnosis of pain syndromes associated to a neuropathic (NP) or somatic component (non-neuropathic pain, NNP). METHODS A study was conducted consisting of two phases: cultural adaptation into the Spanish language by means of conceptual equivalence, including forward and backward translations in duplicate and cognitive debriefing, and testing of psychometric properties in patients with NP (peripheral, central and mixed) and NNP. The analysis of psychometric properties included reliability (internal consistency, inter-rater agreement and test-retest reliability) and validity (ROC curve analysis, agreement with the reference diagnosis and determination of sensitivity, specificity, and positive and negative predictive values in different subsamples according to type of NP). RESULTS A sample of 164 subjects (99 women, 60.4%; age: 60.4 +/- 16.0 years), 94 (57.3%) with NP (36 with peripheral, 32 with central, and 26 with mixed pain) and 70 with NNP was enrolled. The questionnaire was reliable [Cronbach's alpha coefficient: 0.71, inter-rater agreement coefficient: 0.80 (0.71-0.89), and test-retest intra-class correlation coefficient: 0.95 (0.92-0.97)] and valid for a cut-off value > or = 4 points, which was the best value to discriminate between NP and NNP subjects. DISCUSSION This study, representing the first validation of the DN4 questionnaire into another language different than the original, not only supported its high discriminatory value for identification of neuropathic pain, but also provided supplemental psychometric validation (i.e. test-retest reliability, influence of educational level and pain intensity) and showed its validity in mixed pain syndromes.
Resumo:
Background: Current guidelines underline the limitations of existing instruments to assess fitness to drive and the poor adaptability of batteries of neuropsychological tests in primary care settings. Aims: To provide a free, reliable, transparent computer based instrument capable of detecting effects of age or drugs on visual processing and cognitive functions. Methods: Relying on systematic reviews of neuropsychological tests and driving performances, we conceived four new computed tasks measuring: visual processing (Task1), movement attention shift (Task2), executive response, alerting and orientation gain (Task3), and spatial memory (Task4). We then planned five studies to test MedDrive's reliability and validity. Study-1 defined instructions and learning functions collecting data from 105 senior drivers attending an automobile club course. Study-2 assessed concurrent validity for detecting minor cognitive impairment (MCI) against useful field of view (UFOV) on 120 new senior drivers. Study-3 collected data from 200 healthy drivers aged 20-90 to model age related normal cognitive decline. Study-4 measured MedDrive's reliability having 21 healthy volunteers repeat tests five times. Study-5 tested MedDrive's responsiveness to alcohol in a randomised, double-blinded, placebo, crossover, dose-response validation trial including 20 young healthy volunteers. Results: Instructions were well understood and accepted by all senior drivers. Measures of visual processing (Task1) showed better performances than the UFOV in detecting MCI (ROC 0.770 vs. 0.620; p=0.048). MedDrive was capable of explaining 43.4% of changes occurring with natural cognitive decline. In young healthy drivers, learning effects became negligible from the third session onwards for all tasks except for dual tasking (ICC=0.769). All measures except alerting and orientation gain were affected by blood alcohol concentrations. Finally, MedDrive was able to explain 29.3% of potential causes of swerving on the driving simulator. Discussion and conclusions: MedDrive reveals improved performances compared to existing computed neuropsychological tasks. It shows promising results both for clinical and research purposes.
Resumo:
Most ventricular assist devices (VADs) currently used in infants are extracorporeal. These VADs require long-term anticoagulation therapy and extensive surgery, and two devices are needed for biventricular support. We designed a biventricular assist device based on shape memory alloy that reproduces the hemodynamic effects of cardiomyoplasty, supporting the heart with a compressing movement, and evaluated its performance in a dedicated mockup system. Nitinol fibers are the device's key component. Ejection fraction (EF), cardiac output (CO), and generated systolic pressure were measured on a test bench. Our test bench settings were a preload range of 0-15 mm Hg, an afterload range of 0-160 mm Hg, and a heart rate (HR) of 20, 30, 40, and 60 beats/min. A power supply of 15 volts and 3.5 amperes was necessary. The EF range went from 34.4% to 1.2% as the afterload and HR increased, along with a CO from 180 to 6 ml/min. The device generated a maximal systolic pressure of 25 mm Hg. Cardiac compression for biventricular assistance in child-sized heart using shape memory alloy is technically feasible. Further testing remains necessary to assess this VAD's in vivo performance range and its reliability.
Resumo:
Topological indices have been applied to build QSAR models for a set of 20 antimalarial cyclic peroxy cetals. In order to evaluate the reliability of the proposed linear models leave-n-out and Internal Test Sets (ITS) approaches have been considered. The proposed procedure resulted in a robust and consensued prediction equation and here it is shown why it is superior to the employed standard cross-validation algorithms involving multilinear regression models
Resumo:
The purpose of this paper is to describe the development and to test the reliability of a new method called INTERMED, for health service needs assessment. The INTERMED integrates the biopsychosocial aspects of disease and the relationship between patient and health care system in a comprehensive scheme and reflects an operationalized conceptual approach to case mix or case complexity. The method is developed to enhance interdisciplinary communication between (para-) medical specialists and to provide a method to describe case complexity for clinical, scientific, and educational purposes. First, a feasibility study (N = 21 patients) was conducted which included double scoring and discussion of the results. This led to a version of the instrument on which two interrater reliability studies were performed. In study 1, the INTERMED was double scored for 14 patients admitted to an internal ward by a psychiatrist and an internist on the basis of a joint interview conducted by both. In study 2, on the basis of medical charts, two clinicians separately double scored the INTERMED in 16 patients referred to the outpatient psychiatric consultation service. Averaged over both studies, in 94.2% of all ratings there was no important difference between the raters (more than 1 point difference). As a research interview, it takes about 20 minutes; as part of the whole process of history taking it takes about 15 minutes. In both studies, improvements were suggested by the results. Analyses of study 1 revealed that on most items there was considerable agreement; some items were improved. Also, the reference point for the prognoses was changed so that it reflected both short- and long-term prognoses. Analyses of study 2 showed that in this setting, less agreement between the raters was obtained due to the fact that the raters were less experienced and the scoring procedure was more susceptible to differences. Some improvements--mainly of the anchor points--were specified which may further enhance interrater reliability. The INTERMED proves to be a reliable method for classifying patients' care needs, especially when used by experienced raters scoring by patient interview. It can be a useful tool in assessing patients' care needs, as well as the level of needed adjustment between general and mental health service delivery. The INTERMED is easily applicable in the clinical setting at low time-costs.
Resumo:
Objective: To test the efficacy of teaching motivational interviewing (MI) to medical students. Methods: Thirteen 4th year medical students volunteered to participate. Seven days before and 7 days after an 8-hour interactive training MI workshop, each student performed a videorecorded interview with two standardized patients: a 60 year old alcohol dependent woman and a 50 year old cigarette smoking man. Students' counseling skills were coded by two blinded clinicians using the Motivational Interviewing Treatment Integrity 3.0 (MITI). Inter-rater reliability was calculated for all interviews and a test-retest was completed in a sub-sample of 10 consecutive interviews three days apart. Difference between MITI scores before and after training were calculated and tested using non-parametric tests. Effect size was approximated by calculating the probability that posttest scores are greater than pretest scores (P*=P(Pre<Post)+1/2P(Pre=Post)), P*>1/2 indicating greater scores in posttest, P*=1/2 no effect, and P*<1/2 smaller scores in posttest. Results: Median differences between MITI scores before and after MI training indicated a general progression in MI skills: MI spirit global score (median difference=1.5, Inter quartile range=1.5, p<0.001, P*=0.90); Empathy global score (med diff=1, IQR=0.5, p<0.001, P*=0.85); Percentage of MI adherent skills (med diff=36.6, IQR=50.5, p<0.001, P*=0.85); Percentage of open questions (med diff=18.6, IQR=21.6, p<0.001, P*=0.96); reflections/ questions ratio (med diff=0.2, IQR=0.4, p<0.001, P*=0.81). Only Direction global score and the percentage of complex reflections were not significantly improved (med diff=0, IQR=1, p=0.53, P*=0.44, and med diff=4.3, IQR=24.8, p=0.48, P*=0.62, respectively). Inter-rater reliability indicated weighted kappa ranged between 0.14 for Direction to 0.51 for Collaboration and ICC ranged between 0.28 for Simple reflection to 0.95 for Closed question. Test-retests indicated weighted kappa ranged between 0.27 for Direction to 0.80 for Empathy and ICC ranged between 0.87 for Complex reflection to 0.98 for Closed question. Conclusion: This pilot study indicated that an 8-hour training in MI for voluntary 4th year medical students resulted in significant improvement of MI skills. Larger sample of unselected medical students should be studied to generalize the benefit of MI training to medical students. Interrater reliability and test-retests suggested that coders' training should be intensified.
Resumo:
PURPOSE: To conduct a cross-cultural adaptation of the Core Outcome Measures Index (COMI) into French according to established guidelines. METHODS: Seventy outpatients with chronic low back pain were recruited from six spine centres in Switzerland and France. They completed the newly translated COMI, and the Roland Morris disability (RMQ), Dallas Pain (DPQ), adjectival pain rating scale, WHO Quality of Life, and EuroQoL-5D questionnaires. After ~14 days RMQ and COMI were completed again to assess reproducibility; a transition question (7-point Likert scale; "very much worse" through "no change" to "very much better") indicated any change in status since the first questionnaire. RESULTS: COMI whole scores displayed no floor effects and just 1.5% ceiling effects. The scores for the individual COMI items correlated with their corresponding full-length reference questionnaire with varying strengths of correlation (0.33-0.84, P < 0.05). COMI whole scores showed a very good correlation with the "multidimensional" DPQ global score (Rho = 0.71). 55 patients (79%) returned a second questionnaire with no/minimal change in their back status. The reproducibility of individual COMI 5-point items was good, with test-retest differences within one grade ranging from 89% for 'social/work disability' to 98% for 'symptom-specific well-being'. The intraclass correlation coefficient for the COMI whole score was 0.85 (95% CI 0.76-0.91). CONCLUSIONS: In conclusion, the French version of this short, multidimensional questionnaire showed good psychometric properties, comparable to those reported for German and Spanish versions. The French COMI represents a valuable tool for future multicentre clinical studies and surgical registries (e.g. SSE Spine Tango) in French-speaking countries.
Resumo:
When facing age-related cerebral decline, older adults are unequally affected by cognitive impairment without us knowing why. To explore underlying mechanisms and find possible solutions to maintain life-space mobility, there is a need for a standardized behavioral test that relates to behaviors in natural environments. The aim of the project described in this paper was therefore to provide a free, reliable, transparent, computer-based instrument capable of detecting age-related changes on visual processing and cortical functions for the purposes of research into human behavior in computational transportation science. After obtaining content validity, exploring psychometric properties of the developed tasks, we derived (Study 1) the scoring method for measuring cerebral decline on 106 older drivers aged ≥70 years attending a driving refresher course organized by the Swiss Automobile Association to test the instrument's validity against on-road driving performance (106 older drivers). We then validated the derived method on a new sample of 182 drivers (Study 2). We then measured the instrument's reliability having 17 healthy, young volunteers repeat all tests included in the instrument five times (Study 3) and explored the instrument's psychophysical underlying functions on 47 older drivers (Study 4). Finally, we tested the instrument's responsiveness to alcohol and effects on performance on a driving simulator in a randomized, double-blinded, placebo, crossover, dose-response, validation trial including 20 healthy, young volunteers (Study 5). The developed instrument revealed good psychometric properties related to processing speed. It was reliable (ICC = 0.853) and showed reasonable association to driving performance (R (2) = 0.053), and responded to blood alcohol concentrations of 0.5 g/L (p = 0.008). Our results suggest that MedDrive is capable of detecting age-related changes that affect processing speed. These changes nevertheless do not necessarily affect driving behavior.
Resumo:
Objective: To evaluate the internal consistency of the version of the Michigan Alcoholism Screening Test – Geriatric Version (MAST-G) instrument, translated and adapted for Brazil. Method: This was a descriptive, cross-sectional study. Data were collected through a demographic questionnaire, the ICD-10 and the MAST-G, following the steps of translation and cultural adaptation. One hundred eleven elderly in the city of São Carlos, SP, Brazil were interviewed. Results: The mean age of those interviewed was 70 years, with 45% men and 55% women, with the mean education of three years; 92% resided with family; 48% of the subjects consumed alcoholic beverages. The MAST-G presented a good level of reliability, with Cronbach’s α = 0.7873, and good levels of sensitivity and specificity with a cutoff score of five positive responses. Conclusion: The Brazilian version of the MAST-G presented internal consistency values similar to the original English version,showing it to be adequate for use in the national context.
Resumo:
This paper proposes a nonparametric test in order to establish the level of accuracy of theforeign trade statistics of 17 Latin American countries when contrasted with the trade statistics of the main partners in 1925. The Wilcoxon Matched-Pairs Ranks test is used to determine whether the differences between the data registered by exporters and importers are meaningful, and if so, whether the differences are systematic in any direction. The paper tests for the reliability of the data registered for two homogeneous products, petroleum and coal, both in volume and value. The conclusion of the several exercises performed is that we cannot accept the existence of statistically significant differences between the data provided by the exporters and the registered by the importing countries in most cases. The qualitative historiography of Latin American describes its foreign trade statistics as mostly unusable. Our quantitative results contest this view.
Resumo:
Purpose: Many countries used the PGMI (P=perfect, G=good, M=moderate, I=inadequate) classification system for assessing the quality of mammograms. Limits inherent to the subjectivity of this classification have been shown. Prior to introducing this system in Switzerland, we wanted to better understand the origin of this subjectivity in order to minimize it. Our study aimed at identifying the main determinants of the variability of the PGMI system and which criteria are the most subjected to subjectivity. Methods and Materials: A focus group composed of 2 experienced radiographers and 2 radiologists specified each PGMI criterion. Ten raters (6 radiographers and 4 radiologists) evaluated twice a panel of 40 randomly selected mammograms (20 analogic and 20 digital) according to these specified PGMI criteria. The PGMI classification was assessed and the intra- and inter-rater reliability was tested for each professional group (radiographer vs radiologist), image technology (analogic vs digital) and PGMI criterion. Results: Some 3,200 images were assessed. The intra-rater reliability appears to be weak, particularly in respect to inter-rater variability. Subjectivity appears to be largely independent of the professional group and image technology. Aspects of the PGMI classification criteria most subjected to variability were identified. Conclusion: Post-test discussions enabled to specify more precisely some criteria. This should reduce subjectivity when applying the PGMI classification system. A concomitant, important effort in training radiographers is also necessary.
Resumo:
Reducing a test administration to standardised procedures reflects the test designers' standpoint. However, from the practitioners' standpoint, each client is unique. How do psychologists deal with both standardised test administration and clients' diversity? To answer this question, we interviewed 17 psychologists working in three public services for children and adolescents about their assessment practices. We analysed the numerous "client categorisations" they produced in their accounts. We found that they had shared perceptions about their clients' diversity, and reported various non-standard practices that complemented standardised test administration, but also differed from them or were even forbidden. They seem to experience a dilemma between: (a) prescribed and situated practices; (b) scientific and situated reliability; (c) commutative and distributive justice. For practitioners, dealing with clients' diversity this is a practical problem, halfway between a problem-solving task and a moral dilemma.
Resumo:
Introduction: Carbon monoxide (CO) poisoning is one of the mostcommon causes of fatal poisoning. Symptoms of CO poisoning arenonspecific and the documentation of elevated carboxyhemoglobin(HbCO) levels in arterial blood sample is the only standard ofconfirming suspected exposure. The treatment of CO poisoning requiresnormobaric or hyperbaric oxygen therapy, according to the symptomsand HbCO levels. A new device, the Rad-57 pulse CO-oximeter allowsnoninvasive transcutaneous measurement of blood carboxyhemoglobinlevel (SpCO) by measurement of light wavelength absorptions.Methods: Prospective cohort study with a sample of patients, admittedbetween October 2008 - March 2009 and October 2009 - March 2010,in the emergency services (ES) of a Swiss regional hospital and aSwiss university hospital (Burn Center). In case of suspected COpoisoning, three successive noninvasive measurements wereperformed, simultaneously with one arterial blood HbCO test. A controlgroup includes patients admitted in the ES for other complaints (cardiacinsufficiency, respiratory distress, acute renal failure), but necessitatingarterial blood testing. Informed consent was obtained from all patients.The primary endpoint was to assess the agreement of themeasurements made by the Rad-57 (SpCO) and the blood levels(HbCO).Results: 50 patients were enrolled, among whom 32 were admittedfor suspected CO poisoning. Baseline demographic and clinicalcharacteristics of patients are presented in table 1. The median age was37.7 ans ± 11.8, 56% being male. Median laboratory carboxyhemoglobinlevels (HbCO) were 4.25% (95% IC 0.6-28.5) for intoxicated patientsand 1.8% (95% IC 1.0-5.3) for control patients. Only five patientspresented with HbCO levels >= 15%. The results disclose relatively faircorrelations between the SpCO levels obtained by the Rad-57 and thestandard HbCO, without any false negative results. However, theRad-57 tend to under-estimate the value of SpCO for patientsintoxicated HbCO levels >10% (fig. 1).Conclusion: Noninvasive transcutaneous measurement of bloodcarboxyhemoglobin level is easy to use. The correlation seems to becorrect for low to moderate levels (<15%). For higher values, weobserve a trend of the Rad-57 to under-estimate the HbCO levels. Apartfrom this potential limitation and a few cases of false-negative resultsdescribed in the literature, the Rad-57 may be useful for initial triageand diagnosis of CO.