This is the second of a series of articles about how institutions attempt to measure results.

* * * * *

What the government found when it measured itself

On the ninth of February, 1968, at the Conrad Hilton in Chicago, Alice Rivlin told a room of education researchers what two years of the best analytical method in American government had produced.

She was thirty-six years old, Deputy Assistant Secretary of Health, Education, and Welfare for Program Analysis, one of the two deputies in the department’s analytical office. After two years of applying the Planning-Programming-Budgeting System to educational programs at HEW, she said, they had only systematized their ignorance. They had a much clearer idea than before of just what it was they did not know.

I want to be exact about how that sentence reaches us, because the manner of its survival belongs to the subject. The talk was never published. It is not in the education research archive and it is not on her own list of papers. The typescript is held at the Library of Congress, offsite and undigitized. The most quoted line in the literature of federal program budgeting comes down to us because a researcher named Roger Kaufman printed it in a monograph two years later, and nowhere else.

She said our ignorance. She was not reporting on somebody else’s failure.

The instruction that produced that result had gone out twenty-eight months before.

*   *   *

The previous essay, The Body Count, ended on it. Seven weeks after Johnson announced the system at a Cabinet breakfast, the Bureau of the Budget put Bulletin 66-3 in the mail — twenty-two departments required to comply, seventeen more encouraged, ten days to name a responsible official.

Buried on the fourth page was the instruction that mattered. Agencies were to express their objectives and planned accomplishments, wherever possible, in quantitative non-financial terms. Youths trained. Broadcast hours. Children. Patients.

Five months later two men at the State Department wrote a memorandum that named the problem and then argued past it.

Walt Rostow and William Crockett sent it up to Secretary Rusk on the seventh of March, 1966. They were not resisting. They praised the Pentagon approach as tested and successful, and what they wanted was a programming system of their own, built around countries rather than around the Bureau of the Budget’s uniform template. In the middle of that argument they wrote the sentence.

Cost benefit analysis in defense relates resources to our ability to kill people. Foreign affairs, they went on, offered no simple mathematical equivalent, and a simple quantitative approach to student exchanges or agricultural credit would be illusory.

That is Part I of this series described in one line, five months in, by two men who admired the thing they were describing.

*   *   *

Rivlin had come to HEW the month before.

She believed measurement could improve government and had come to Washington to do it. She also knew what the men at Defense had that she did not. A bomber sortie is an output. A child’s education is not, and no amount of goodwill converts one into the other.

That same year Allen Schick named the thing the whole reform was resting on. Writing in Public Administration Review in December 1966, he set out the assumption in one clause — that behavior follows form, that changing how information is classified changes what people do with it. Take the assumption away, he wrote, and the movement is reduced to manipulating techniques.

Allen Schick Article Public Administration Review

He did not say the assumption was false.

He said it was load-bearing, and that nobody had examined it.

Seven years later he came back and said what it had cost.

*   *   *

In the meantime, the numbers went to work.

The Office of Economic Opportunity had an analytical shop under Joseph Kershaw and then Robert Levine, doing the kind of work the bulletin described. Levine published his figures in 1966: Job Corps at a steady-state cost per graduate of $6,980, and a cost per success of $8,725, against roughly a thousand and twenty-two hundred for the out-of-school Neighborhood Youth Corps. He called the calculation highly hypothetical and said it rested on very sketchy data.

That autumn the House capped Job Corps costs by statute at seventy-five hundred dollars per enrollee, on an amendment offered by Edith Green of Oregon after a five-thousand-dollar version had failed by three votes. Thirteen months separate Bulletin 66-3 from that vote. The debate has been searched. There is no mention of program budgeting, no systems analysis, no reference to the bulletin. Sar Levitan, writing the following year, attributed the ceiling to public criticism of Job Corps costs and to sniping from officials of competing federal programs.

What the ceiling did establish was where the argument would now be held. In the first ten months of fiscal 1967 the total annual cost per enrollee ran about eighty-one hundred dollars, above the limit and lawful, because the statute capped direct operating costs at centers open more than nine months and the totals carried overhead, capital amortization, and the materials enrollees used on public projects.

Measured the way the law measured, Job Corps was under the limit.

*   *   *

The man who ran the analytical shop said the boundary out loud the following year.

William Gorham testified on the fourteenth of September, 1967, before the Joint Economic Committee’s Subcommittee on Economy in Government. They had attempted no grand cost-benefit analysis of health against education against welfare, he said, and if I was ever naive enough to think this sort of analysis possible — he no longer was. No amount of analysis would tell the nation whether it benefits more from sending a slum child to preschool, providing medical care to an old man, or enabling a disabled housewife to resume her normal activities. Those were questions of value and politics, and on them the analyst could not make much contribution.

The order he was describing is old, though the man who wrote it down was arguing the opposite case. Federalist 62 puts good government at two things in sequence — fidelity to the object, which is the happiness of the people, and then, second, knowledge of the means by which that object may best be attained. Madison’s complaint was that America had neglected the second, and his remedy was a Senate term long enough to acquire it. He was not warning anyone off the means. He was fixing the order they come in. Gorham’s shop was built for the second and had just reported that it could not reach the first.

*   *   *

Before this becomes a case against measuring government, the other side has to win an argument, and it wins more of it than the story so far suggests.

In 2009 Steven Kelman and John Friedman examined the English National Health Service’s four-hour emergency-room target across all one hundred fifty-five hospital trusts in England. Every serious critic had predicted gaming. They took five specific hypotheses — that clinical quality would fall, that resources would be pulled from elective work, that the mean wait would rise behind an improving headline, that a measurement week would show a blip, that patients would be admitted to wards to stop the clock — and tested each. Not one was confirmed. Waits fell sharply.

Four years earlier, Thomas Locker and Suzanne Mason had plotted how long patients actually spent in eighty-three English emergency departments and found the distribution piling up just short of the four-hour mark.

Both are true, and they do not collide. Locker and Mason measured individual departure times inside departments. Kelman and Friedman measured system-level outcomes across trusts, and a pile-up at 240 minutes was not among the five things they tested. A target can bend the clock at the margin and still reduce total waiting without producing any of the failures the theory predicted. That reconciliation is mine, and I have not found an author who states it in print, so take it as a reading rather than a finding.

Government that does not count itself cannot be checked.

The comparison is never the measured system against a perfect one.

*   *   *

Which makes Rivlin’s sentence in Chicago the harder result rather than the easier one.

She published the argument the following year, in a paper for the Joint Economic Committee, and there she put it as a positive finding: the most important result of the effort so far had been the discovery of how little is really known. Anyone who had thought the system a magic formula, she added, had better think again.

That is not a complaint about counting.

It is a report that the counting had worked.

The department could say where the education money went, what it bought, and which children it reached. Every one of those readings was real and none of them answered the question the department existed to answer, and it took two years and a staff of analysts to establish that the answer was not in the file.

The apparatus went on growing. The fiscal 1969 budget carried 1,145 positions for the system across twenty-one agencies, most of them added in the preceding three years. But the Bureau of the Budget’s own accounting is more candid than the total suggests. Of the 825 professionals, only about a third were net additions. The rest came, in the Bureau’s phrase, from revision or rechristening of other jobs.

*   *   *

The end came on the twenty-first of June, 1971, in a memorandum accompanying Circular A-11.

Agencies would no longer submit the multi-year plans, the program memoranda, or the special analytic studies. The memorandum described this as part of a continuing effort to simplify budget submission requirements. Schick, who had the document in front of him, wrote that no mention was made in it of the three initials that had dazzled the world of budgeting five years earlier, and no admission of failure. By these words, he wrote, the thing became an unthing.

In the same article he answered the question he had raised in 1966. Budgeting, he wrote, is the routinization of public choice — standard procedures, timetables, classifications, rules — and PPB failed because it never got inside those routines. The assumption had been that changing the form would change the behavior. The form changed. The routine did not let it in.

It did not die everywhere. What ended was the government-wide requirement. Defense had built the method, Defense kept it, and it runs there still under a name that acquired a fourth term for execution in 2003, governing the budget of the department that spends most of the discretionary dollars. It sat across the river the whole time, fully staffed, which turns out to be a more effective form of invisibility than disappearing.

Rivlin gave her own verdict in the Gaither Lectures, delivered at Berkeley in January 1970 and published the following year. In HEW, she said, the system’s most important effect had been the creation of an analytical staff at the department level — a group of people trained to think analytically whose job was to improve the way decisions got made.

The system answered almost nothing about any program.

It produced the people who would spend the next thirty years asking.

On the twenty-fourth of February, 1975, Alice Rivlin was appointed the first Director of the Congressional Budget Office. She and Robert Reischauer and two assistants worked out of a single shared office in the Dirksen building.

*   *   *

Sixteen years later the Senate went looking for a new system and found one it had already had.

Joseph Wholey testified on the twenty-third of May, 1991, before the Committee on Governmental Affairs, which was drafting what became the Government Performance and Results Act. He had worked on program evaluation at HEW in the years this essay covers. What he said that day cannot be established — no print of the hearing has been located. What can be established is what the committee did with him. He appears exactly once in its report, in the list of witnesses. The report quotes a senator, the General Accounting Office, the Director of the Office of Management and Budget, the Comptroller General, the State of Florida, the city of Sunnyvale, Australia, and the United Kingdom. It never quotes Wholey.

The report named the ancestors in a single sentence pair. Past efforts at comprehensive management reform, it said — PPBS and Zero-Based Budgeting among them — though equally well-intended, had been less than satisfactory. Then: New information technologies, unavailable in past decades, should now be a great advantage.

That is a diagnosis and it does not match the record. Schick had located the failure in 1966 and named it in 1973, and it had nothing to do with computing. The report cites no document from the PPBS years at all. Its borrowings are entirely contemporary.

In July of 1993, six weeks after the report and three weeks before the signing, the Congressional Budget Office published a study requested by the committee’s chairman. It traced the new statute back through performance budgeting, PPBS, and Zero-Based Budgeting, found that each had fallen short of its goals, and then said what each had left behind. These systems may have had an unintended consequence, it wrote — an increased demand for analysis, the capability to do it, and its use on public problems. This, the study said, may be viewed as the most lasting legacy of the “rational budgeting” movement.

That is Rivlin’s 1971 conclusion, reached independently, by the office she had founded, twenty-two years later.

The Government Performance and Results Act was signed on the third of August, 1993. Agencies would set measurable goals and report against them annually. The framework is still law.

*   *   *

The strongest evidence in this essay arrived eighteen years after that, and it is about the program the ceiling was written for.

On the thirty-first of January, 2011, Jane Fortson and Peter Schochet delivered a report to the Employment and Training Administration of the Labor Department. Job Corps had rated its centers for decades, and contractor payments, bonuses, and renewals turned on the rating. Fortson and Schochet set the ratings against impact estimates from the National Job Corps Study, which had assigned applicants at random and could therefore say what a center had actually done to the students who attended it.

Students at highly rated centers did better than students elsewhere. So did the control-group members who would have been assigned to those centers and never attended them.

The ratings were adjusted for student characteristics to see whether that would fix it. It did not. The rankings moved and the adjusted ratings remain uncorrelated with center-level impacts. The Labor Department’s own summary of its own study says that neither the aggregate measure nor any of its components can be reliably associated with impacts on student outcomes.

The report cards are still published.

*   *   *   *   *

The next essay goes to the unit itself. To weigh anything against anything you need a common measure, and where the one this country settled on came from — and what the man who built it printed on page 7 of his own report — is where this goes next.

Charles Cranston Jett is an author, civic educator, and Professional Certified Coach based in Chicago. A graduate of the U.S. Naval Academy (Class of 1964) and Harvard Business School, he served aboard nuclear submarines during the Cold War. He is the author of six books, including Super Nuke!, hosts four podcasts, and writes across his Critical Skills Blog platform on history, leadership, and the health of the American republic. In his writing he uses AI tools for research, editing, and occasional image creation; the arguments, the voice, and the final judgment are his. He and his wife, Dr. Nancy Church, live at Water Tower Residences, where they co-host the Chicago Salons.

Leave a Reply

This site uses Akismet to reduce spam. Learn how your comment data is processed.

Recent Posts

Here are a few of the most recent posts. If you want to search for posts – there are over 1000 articles on this site, simply click on “Search” in the navigation menu at the top. And don’t forget to subscribe. This is a free site and articles are posted frequently at no cost. Please share these articles and leave comments as you deem appropriate.

Discover more from Critical Skills

Subscribe now to keep reading and get access to the full archive.

Continue reading