首页 | 本学科首页   官方微博 | 高级检索  
相似文献
 共查询到20条相似文献,搜索用时 46 毫秒
1.
This paper considers the problem of selecting optimal bandwidths for variable (sample‐point adaptive) kernel density estimation. A data‐driven variable bandwidth selector is proposed, based on the idea of approximating the log‐bandwidth function by a cubic spline. This cubic spline is optimized with respect to a cross‐validation criterion. The proposed method can be interpreted as a selector for either integrated squared error (ISE) or mean integrated squared error (MISE) optimal bandwidths. This leads to reflection upon some of the differences between ISE and MISE as error criteria for variable kernel estimation. Results from simulation studies indicate that the proposed method outperforms a fixed kernel estimator (in terms of ISE) when the target density has a combination of sharp modes and regions of smooth undulation. Moreover, some detailed data analyses suggest that the gains in ISE may understate the improvements in visual appeal obtained using the proposed variable kernel estimator. These numerical studies also show that the proposed estimator outperforms existing variable kernel density estimators implemented using piecewise constant bandwidth functions.  相似文献   

2.
Principal component regression (PCR) has two steps: estimating the principal components and performing the regression using these components. These steps generally are performed sequentially. In PCR, a crucial issue is the selection of the principal components to be included in regression. In this paper, we build a hierarchical probabilistic PCR model with a dynamic component selection procedure. A latent variable is introduced to select promising subsets of components based upon the significance of the relationship between the response variable and principal components in the regression step. We illustrate this model using real and simulated examples. The simulations demonstrate that our approach outperforms some existing methods in terms of root mean squared error of the regression coefficient.  相似文献   

3.
Least-squares regression is not appropriate when the response variable is circular, and can lead to erroneous results. The reason for this is that the squared difference is not an appropriate measure of distance on the circle. In this paper, a circular analog to least-squares regression is presented for predicting a circular response variable by another circular variable and a set of linear covariates. An alternative maximum-likelihood formulation yields the same regression parameter estimates. Under the maximum-likelihood model, asymptotic standard errors of the parameter estimates are obtained. As an example, the regression model is used to model data from a marine biology study.  相似文献   

4.
A new method is proposed for measuring the distance between a training data set and a single, new observation. The novel distance measure reflects the expected squared prediction error when a quantitative response variable is predicted on the basis of the training data set using the distance weighted k-nearest-neighbor method. The simulation presented here shows that the distance measure correlates well with the true expected squared prediction error in practice. The distance measure can be applied, for example, in assessing the uncertainty of prediction.  相似文献   

5.
In this paper, we introduce a new partially functional linear varying coefficient model, where the response is a scalar and some of the covariates are functional. By means of functional principal components analysis and local linear smoothing techniques, we obtain the estimators of coefficient functions of both function-valued variable and real-valued variables. Then the rates of convergence of the proposed estimators and the mean squared prediction error are established under some regularity conditions. Moreover, we develop a hypothesis test for the model and employ the bootstrap procedure to evaluate the null distribution of test statistic and the p-value of the test. At last, we illustrate the finite sample performance of our methods with some simulation studies and a real data application.  相似文献   

6.
The measurable multiple bio-markers for a disease are used as indicators for studying the response variable of interest in order to monitor and model disease progression. However, it is common for subjects to drop out of the studies prematurely resulting in unbalanced data and hence complicating the inferences involving such data. In this paper we consider a case where data are unbalanced among subjects and also within a subject because for some reason only a subset of the multiple outcomes of the response variable are observed at any one occasion. We propose a nonlinear mixed-effects model for the multivariate response variable data and derive a joint likelihood function that takes into account the partial dropout of the outcomes of the response variable. We further show how the methodology can be used in the estimation of the parameters that characterise HIV disease dynamics. An approximation technique of the parameters is also given and illustrated using a routine observational HIV dataset.  相似文献   

7.
In this paper, we investigate the relationship between a functional random covariable and a scalar response which is subject to left-truncation by another random variable. Precisely, we use the mean squared relative error as a loss function to construct a nonparametric estimator of the regression operator of these functional truncated data. Under some standard assumptions in functional data analysis, we establish the almost sure consistency, with rates, of the constructed estimator as well as its asymptotic normality. Then, a simulation study, on finite-sized samples, was carried out in order to show the efficiency of our estimation procedure and to highlight its superiority over the classical kernel estimation, for different levels of simulated truncated data.  相似文献   

8.
In this paper, two new multiple influential observation detection methods, GCD.GSPR and mCD*, are introduced for logistic regression. The proposed diagnostic measures are compared with the generalized difference in fits (GDFFITS) and the generalized squared difference in beta (GSDFBETA), which are multiple influential diagnostics. The simulation study is conducted with one, two and five independent variable logistic regression models. The performance of the diagnostic measures is examined for a single contaminated independent variable for each model and in the case where all the independent variables are contaminated with certain contamination rates and intensity. In addition, the performance of the diagnostic measures is compared in terms of the correct identification rate and swamping rate via a frequently referred to data set in the literature.  相似文献   

9.
The lasso procedure is an estimator‐shrinkage and variable selection method. This paper shows that there always exists an interval of tuning parameter values such that the corresponding mean squared prediction error for the lasso estimator is smaller than for the ordinary least squares estimator. For an estimator satisfying some condition such as unbiasedness, the paper defines the corresponding generalized lasso estimator. Its mean squared prediction error is shown to be smaller than that of the estimator for values of the tuning parameter in some interval. This implies that all unbiased estimators are not admissible. Simulation results for five models support the theoretical results.  相似文献   

10.
Whenever there is auxiliary information available in any form, the researchers want to utilize it in the method of estimation to obtain the most efficient estimator. When there exists enough amount of correlation between the study and the auxiliary variables, and parallel to these associations, the ranks of the auxiliary variables are also correlated with the study variable, which can be used a valuable device for enhancing the precision of an estimator accordingly. This article addresses the problem of estimating the finite population mean that utilizes the complementary information in the presence of (i) the auxiliary variable and (ii) the ranks of the auxiliary variable for non response. We suggest an improved estimator for estimating the finite population mean using the auxiliary information in the presence of non response. Expressions for bias and mean squared error of considered estimators are derived up to the first order of approximation. The performance of estimators is compared theoretically and numerically. A numerical study is carried out to evaluate the performances of estimators. It is observed that the proposed estimator is more efficient than the usual sample mean and the regression estimators, and some other families of ratio and exponential type of estimators.  相似文献   

11.
Expectile regression [Newey W, Powell J. Asymmetric least squares estimation and testing, Econometrica. 1987;55:819–847] is a nice tool for estimating the conditional expectiles of a response variable given a set of covariates. Expectile regression at 50% level is the classical conditional mean regression. In many real applications having multiple expectiles at different levels provides a more complete picture of the conditional distribution of the response variable. Multiple linear expectile regression model has been well studied [Newey W, Powell J. Asymmetric least squares estimation and testing, Econometrica. 1987;55:819–847; Efron B. Regression percentiles using asymmetric squared error loss, Stat Sin. 1991;1(93):125.], but it can be too restrictive for many real applications. In this paper, we derive a regression tree-based gradient boosting estimator for nonparametric multiple expectile regression. The new estimator, referred to as ER-Boost, is implemented in an R package erboost publicly available at http://cran.r-project.org/web/packages/erboost/index.html. We use two homoscedastic/heteroscedastic random-function-generator models in simulation to show the high predictive accuracy of ER-Boost. As an application, we apply ER-Boost to analyse North Carolina County crime data. From the nonparametric expectile regression analysis of this dataset, we draw several interesting conclusions that are consistent with the previous study using the economic model of crime. This real data example also provides a good demonstration of some nice features of ER-Boost, such as its ability to handle different types of covariates and its model interpretation tools.  相似文献   

12.
Sousa et al. and Gupta et al. suggested ratio and regression-type estimators of the mean of a sensitive variable using nonsensitive auxiliary variable. This article proposes exponential-type estimators using one and two auxiliary variables to improve the efficiency of mean estimator based on a randomized response technique. The expressions for the mean squared errors (MSEs) and bias, up to first-order approximation, have been obtained. It is shown that the proposed exponential-type estimators are more efficient than the existing estimators. The gain in efficiency over the existing estimators has also been shown with a simulation study and by using real data.  相似文献   

13.
Kupper and Meydrech and Myers and Lahoda introduced the mean squared error (MSE) approach to study response surface designs, Duncan and DeGroot derived a criterion for optimality of linear experimental designs based on minimum mean squared error. However, minimization of the MSE of an estimator maxr renuire some knowledge about the unknown parameters. Without such knowledge construction of designs optimal in the sense of MSE may not be possible. In this article a simple method of selecting the levels of regressor variables suitable for estimating some functions of the parameters of a lognormal regression model is developed using a criterion for optimality based on the variance of an estimator. For some special parametric functions, the criterion used here is equivalent to the criterion of minimizing the mean squared error. It is found that the maximum likelihood estimators of a class of parametric functions can be improved substantially (in the sense of MSE) by proper choice of the values of regressor variables. Moreover, our approach is applicable to analysis of variance as well as regression designs.  相似文献   

14.
The mean squared error (MSE)-minimizing local variable bandwidth for the univariate local linear estimator (the LL) is well-known. This bandwidth does not stabilize variance over the domain. Moreover, in regions where a regression function has zero curvature, the LL estimator is discontinuous. In this paper, we propose a variance-stabilizing (VS) local variable diagonal bandwidth matrix for the multivariate LL estimator. Theoretically, the VS bandwidth can outperform the multivariate extension of the MSE-minimizing local variable scalar bandwidth in terms of asymptotic mean integrated squared error and can avoid discontinuity created by the MSE-minimizing bandwidth. We present an algorithm for estimating the VS bandwidth and simulation studies.  相似文献   

15.
Abstract

The availability of some extra information, along with the actual variable of interest, may be easily accessible in different practical situations. A sensible use of the additional source may help to improve the properties of statistical techniques. In this study, we focus on the estimators for calibration and intend to propose a setup where we reply only on first two moments instead of modeling the whole distributional shape. We have proposed an estimator for linear calibration problems and investigated it under normal and skewed environments. We have partitioned its mean squared error into intrinsic and estimation components. We have observed that the bias and mean squared error of the proposed estimator are function of four dimensionless quantities. It is to be noticed that both the classical and the inverse estimators become the special cases of the proposed estimator. Moreover, the mean squared error of the proposed estimator and the exact mean squared error of the inverse estimator coincide. We have also observed that the proposed estimator performs quite well for skewed errors as well. The real data applications are also included in the study for practical considerations.  相似文献   

16.
Log-normal linear models are widely used in applications, and many times it is of interest to predict the response variable or to estimate the mean of the response variable at the original scale for a new set of covariate values. In this paper we consider the problem of efficient estimation of the conditional mean of the response variable at the original scale for log-normal linear models. Several existing estimators are reviewed first, including the maximum likelihood (ML) estimator, the restricted ML (REML) estimator, the uniformly minimum variance unbiased (UMVU) estimator, and a bias-corrected REML estimator. We then propose two estimators that minimize the asymptotic mean squared error and the asymptotic bias, respectively. A parametric bootstrap procedure is also described to obtain confidence intervals for the proposed estimators. Both the new estimators and the bootstrap procedure are very easy to implement. Comparisons of the estimators using simulation studies suggest that our estimators perform better than the existing ones, and the bootstrap procedure yields confidence intervals with good coverage properties. A real application of estimating the mean sediment discharge is used to illustrate the methodology.  相似文献   

17.
In this article, a new class of variance function estimators is proposed in the setting of heteroscedastic nonparametric regression models. To obtain a variance function estimator, the main proposal is to smooth the product of the response variable and residuals as opposed to the squared residuals. The asymptotic properties of the proposed methodology are investigated in order to compare its asymptotic behavior with that of the existing methods. The finite sample performance of the proposed estimator is studied through simulation studies. The effect of the curvature of the mean function on its finite sample behavior is also discussed.  相似文献   

18.
In this article, we introduce and study local constant and local linear nonparametric regression estimators when it is appropriate to assess performance in terms of mean squared relative error of prediction. We give asymptotic results for both boundary and non-boundary cases. These are special cases of more general asymptotic results that we provide concerning the estimation of the ratio of conditional expectations of two functions of the response variable. We also provide a good bandwidth selection method for the estimators. Examples of application, limited simulation results and discussion of related problems and approaches are also given.  相似文献   

19.
The least squares estimation of the slope parameter of a simple linear regression is biased if the regressor variable is measured with random errors. This bias as well as the mean squared error is computed up to the order of 1/T without assuming normality for the error variable. They depend on the fourth moment of the error variable.  相似文献   

20.
In the case where non-experimental data are available from an industrial process and a directed graph for how various factors affect a response variable is known based on a substantive understanding of the process, we consider a problem in which a control plan involving multiple treatment variables is conducted in order to bring a response variable close to a target value with variation reduction. Using statistical causal analysis with linear (recursive and non-recursive) structural equation models, we configure an optimal control plan involving multiple treatment variables through causal parameters. Based on the formulation, we clarify the causal mechanism for how the variance of a response variable changes when the control plan is conducted. The results enable us to evaluate the effect of a control plan on the variance of a response variable from non-experimental data and provide a new application of linear structural equation models to engineering science.  相似文献   

设为首页 | 免责声明 | 关于勤云 | 加入收藏

Copyright©北京勤云科技发展有限公司  京ICP备09084417号