munotes®

Size Metrics: Lines of Code and Function Points

Get access to whole semester resourcesSemester Pass

Chapter Sixty-Seven

Syllabus topic Module 2, "Software Metrics: ... different types of metrics"

Pages 384 to 390 of 622

In one line

Size is the measure most other software measures are divided by, and there are two families of it: lines of code, which count what programmers wrote and mean nothing until the counting rule is stated, and function points, which count what the software does for its users and can be counted before any code exists.

In the wording a student can write in an examination: SLOC, "source lines of code, the number of lines of programming language code in a program before compilation" (ISO/IEC 20968), is counted as physical source lines or logical source statements (Park, SEI 1992), and only a written counting rule makes the number meaningful. A function point is a "unit of measure for functional size", and functional size is the "size of the software derived by quantifying the functional user requirements" (ISO/IEC 20926 and related standards). Function point analysis counts five kinds of component: external inputs (EI), external outputs (EO), external inquiries (EQ), internal logical files (ILF) and external interface files (EIF); rates each low, average or high from its data element types (DET), record element types (RET) or file types referenced (FTR); adds their weights to give the unadjusted function points (UFP); and multiplies by a value adjustment factor, VAF = (TDI × 0.01) + 0.65, where TDI totals the ratings of 14 general system characteristics.

Why size

Park's report opens on why size matters: "Size measures have direct application to the planning, tracking, and estimating of software projects. They are used also to compute productivities, to normalize quality indicators, and to derive measures for memory utilization and test coverage." Normalising is the use testers meet first. Release 2.0 of ExamReg had 200 defects. Whether that is many or few depends on how big release 2.0 is, and a defect density, defects per thousand lines or per function point, is only as trustworthy as the size it divides by.

Lines of code, and the counting problem

Counting lines sounds like the one measure nobody could get wrong. Park's report, written for the SEI's Software Metrics Definition Working Group, found otherwise: "reported values for software size are often confusing and easily misinterpreted. This usually happens because neither the conveyors nor the receivers of the information know what the measurements include or whether the measures have been applied with any consistency." Its example: "reports like 'Our software activity produced 163,000 source code instructions on that job' can easily be misunderstood by a factor of three or more."

The report separates two measures:

  • Physical source lines: lines of the source file, counted by a stated rule about which kinds of line count.
  • Logical source statements: the statements of the language, however many lines each occupies.
munotes.in384

Size Metrics: Lines of Code and Function Points

Its answer to the ambiguity is a definition checklist, on which an organisation ticks, attribute by attribute, what its count includes and excludes. Its basic definition of physical source lines includes executable lines, declarations and compiler directives, and excludes comments (on their own lines or beside code), banners, blank comments and blank lines; and "When a line or statement contains more than one type, classify it as the type with the highest precedence", so a line of code with a comment after it counts as code.

Worked example 1: one module, five sizes

Here is a small module of ExamReg, with a docstring, comments, blank lines and one statement spread over several lines:

"""ExamReg: fees for the examination form (release 2.0).

The fee rules are the exam cell's; the numbers are rupees.
"""

FORM_FEE = 800          # waived with a concession
BACKLOG_FEE = 150       # for each backlog paper

LATE_FEES = {
    "on time": 0,
    "1 to 7 days": 100,
    "8 to 15 days": 500,
}


def late_band(days_late):
    # which row of LATE_FEES applies
    if days_late == 0:
        return "on time"
    if days_late <= 7:
        return "1 to 7 days"
    return "8 to 15 days"


def total_fee(days_late, backlog_papers, concession):
    """Form fee (waived with a concession), backlog fees and the late fee."""
    if days_late < 0 or days_late > 15:
        raise ValueError("form not accepted")
    fee = (0 if concession else FORM_FEE) + BACKLOG_FEE * backlog_papers
    return fee + LATE_FEES[late_band(days_late)]

The program counts it by five rules, each of which someone could honestly call "lines of code".

import ast

source = open("examreg_fees.py").read()
lines = source.splitlines()
tree = ast.parse(source)

docstring_lines = set()                     # lines inside a module or function docstring
for node in [tree] + [n for n in ast.walk(tree) if isinstance(n, ast.FunctionDef)]:
    first = node.body[0]
    if isinstance(first, ast.Expr) and isinstance(first.value, ast.Constant):
        docstring_lines.update(range(first.lineno, first.end_lineno + 1))

every = set(range(1, len(lines) + 1))
nonblank = {n for n in every if lines[n - 1].strip()}
comment = {n for n in nonblank if lines[n - 1].strip().startswith("#")}
statements = [n for n in ast.walk(tree) if isinstance(n, ast.stmt)
              and not (isinstance(n, ast.Expr) and isinstance(n.value, ast.Constant))]

counts = {
    "every line in the file": len(every),
    "nonblank lines": len(nonblank),
    "nonblank, noncomment lines (Park's basic SLOC)": len(nonblank - comment),
    "the same, docstrings treated as comments": len(nonblank - comment - docstring_lines),
    "logical statements (Python statements)": len(statements),
}
for rule, n in counts.items():
    print(f"{rule:<48} {n:>3}")
print(f"largest count / smallest: {max(counts.values()) / min(counts.values()):.1f}")
every line in the file                            30
nonblank lines                                    23
nonblank, noncomment lines (Park's basic SLOC)    22
the same, docstrings treated as comments          18
logical statements (Python statements)            14
largest count / smallest: 2.1

One file of thirty lines has five sizes, from 30 down to 14, a spread of more than two to one on a module chosen to be ordinary. Each rule is defensible, and a report that says only 30 lines of code or 14 lines of code tells the reader nothing about which was used. The row for docstrings shows why the checklist must be filled in for each language: Park's checklist has a row for comments, but a Python docstring is a string the language executes, which the reader thinks of as a comment. Which way to count it is a decision, and it has to be written down.

munotes.in385

Size Metrics: Lines of Code and Function Points

Park's recommendation, after weighing the two measures, is to "adopt physical source lines of code (SLOC) as one of their first measures of software size". Among his reasons: "It is easier and cheaper to build automated counters for physical lines than it is for logical statements", "Rules for determining when logical source statements begin and end are complex and different for every source language", and "Most historical data is in terms of physical source lines."

Function points

Lines of code have two weaknesses no counting rule can cure: they depend on the language (the same fee rule is a different number of lines in Python, Java and COBOL), and they exist only after the code is written, which is too late for estimating the project. Function points measure something else: what the software does for its users. They were defined in 1979 by Allan J. Albrecht at IBM, and the method in use is the International Function Point Users Group's, standardised as ISO/IEC 20926:2009.

Function point analysis is a "method for measuring functional size" (ISO/IEC 20926). It identifies five kinds of component, defined in that standard:

ComponentDefinition (ISO/IEC 20926)ExamReg examples
External input (EI)"elementary process that processes data or control information sent from outside the boundary"Submit the exam form; pay the fee
External output (EO)"elementary process that sends data or control information outside the application's boundary and includes additional processing logic beyond that of an external inquiry"The hall ticket; the payment receipt
External inquiry (EQ)"elementary process that sends data or control information outside the boundary"View registration status
Internal logical file (ILF)"user-recognizable group of logically related data or control information maintained within the boundary of the application being measured"Student records; exam registrations; payments
External interface file (EIF)a group of data "which is referenced by the application being measured, but which is maintained within the boundary of another application"The college's admission records

An elementary process is the "smallest unit of activity that is meaningful to the user". Each component is then rated by how much data it handles. OMG's Automated Function Points standard gives the definitions, taken from ISO/IEC 20926: a data element type (DET) is "a unique user recognizable, non-repeated attribute that is part of an ILF, EIF, EI or EO"; a record element type (RET) is a "user recognizable sub-group of data element types within a data function"; and a file type referenced (FTR) is a "data function (IFL or EIF) read and/or maintained by a transactional function" (the standard's own typing of ILF).

munotes.in386

Size Metrics: Lines of Code and Function Points

Files are rated from their DETs and RETs, transactions from their DETs and FTRs, each rating low, average or high, and each rating carries a weight:

ComponentLowAverageHigh
Internal logical file71015
External interface file5710
External input346
External output457
External inquiry346

The first four rows are OMG's; the inquiry row is Garmus's, from IFPUG's practice. OMG's standard counts no inquiries separately, for a reason worth knowing: "since automated counting tools cannot distinguish between External Inquiries and External Outputs, all External Inquiries will be included in and counted as External Outputs."

The sum of the weights is the unadjusted function point count. The IFPUG method then adjusts it for the system's general characteristics. Garmus lists the fourteen general system characteristics, from data communication to facilitate change, and gives the rule: "Each characteristic is given a value", the values are summed to the total degree of influence (TDI), and the factor is (TDI * .01) + .65 = VAF, as his slide writes it. The adjusted function points are the unadjusted count "multiplied by the Value Adjustment Factor".

Worked example 2: counting ExamReg

A counter has identified ExamReg's functions and counted each one's DETs and its RETs or FTRs. The program rates them by OMG's rules, adds the two inquiries as the counter rated them (IFPUG's rules for inquiries are in its counting manual, which is not on this book's shelf), applies the adjustment, and uses the result to express release 2.0's defects per function point.

LEVEL = {1: "low", 2: "average", 3: "high"}
WEIGHT = {"ILF": (7, 10, 15), "EIF": (5, 7, 10), "EI": (3, 4, 6), "EO": (4, 5, 7), "EQ": (3, 4, 6)}

def band(n, low_max, mid_max):             # 1, 2 or 3 as OMG AFP 1.0 rescales a count
    return 1 if n <= low_max else 2 if n <= mid_max else 3

def complexity(kind, det, other):
    """OMG AFP 1.0: data functions from DETs and RETs, transactions from DETs and FTRs."""
    if kind in ("ILF", "EIF"):
        total = band(det, 19, 50) + band(other, 1, 5)
    elif kind == "EI":
        total = band(det, 4, 15) + band(other, 1, 2)
    else:                                  # EO
        total = band(det, 5, 19) + band(other, 1, 3)
    return 1 if total <= 3 else 2 if total == 4 else 3

# ExamReg's functions as a counter identified them: (kind, name, DETs, RETs or FTRs)
functions = [("ILF", "Student record", 18, 2), ("ILF", "Exam registration", 24, 3),
             ("ILF", "Payment", 9, 1),
             ("EIF", "College admission records", 22, 1), ("EIF", "University paper catalogue", 8, 1),
             ("EI", "Submit the exam form", 16, 3), ("EI", "Pay the fee", 6, 2),
             ("EI", "Update contact details", 5, 1),
             ("EO", "Hall ticket", 14, 3), ("EO", "Payment receipt", 8, 2),
             ("EO", "Registrations by paper report", 21, 3)]
inquiries = [("View registration status", 1), ("Search the paper catalogue", 1)]   # rated by the counter

ufp = 0
for kind, name, det, other in functions:
    level = complexity(kind, det, other)
    ufp += WEIGHT[kind][level - 1]
    print(f"{kind:<4}{name:<32} DET {det:>2}  {'RET' if kind in ('ILF', 'EIF') else 'FTR'} {other}"
          f"  {LEVEL[level]:<8} {WEIGHT[kind][level - 1]:>2}")
for name, level in inquiries:
    ufp += WEIGHT["EQ"][level - 1]
    print(f"EQ  {name:<32} (rated by the counter)  {LEVEL[level]:<8} {WEIGHT['EQ'][level - 1]:>2}")

gsc = {"Data communication": 4, "Distributed data or processing": 1, "Performance objectives": 3,
       "Heavily used configuration": 2, "Transaction rate": 4, "On-line data entry": 5,
       "End-user efficiency": 4, "On-line update": 4, "Complex processing": 2, "Reusability": 1,
       "Conversion and install ease": 1, "Operational ease": 3, "Multiple-site use": 2,
       "Facilitate change": 3}
tdi = sum(gsc.values())
vaf = tdi * 0.01 + 0.65                    # Garmus 2006: (TDI * .01) + .65 = VAF
print(f"unadjusted function points {ufp}; {len(gsc)} characteristics, TDI {tdi};"
      f" VAF {vaf:.2f}; adjusted function points {ufp * vaf:.1f}")
print(f"release 2.0's 200 defects: {200 / (ufp * vaf):.2f} per function point")
munotes.in387

Size Metrics: Lines of Code and Function Points

ILF Student record                   DET 18  RET 2  low       7
ILF Exam registration                DET 24  RET 3  average  10
ILF Payment                          DET  9  RET 1  low       7
EIF College admission records        DET 22  RET 1  low       5
EIF University paper catalogue       DET  8  RET 1  low       5
EI  Submit the exam form             DET 16  FTR 3  high      6
EI  Pay the fee                      DET  6  FTR 2  average   4
EI  Update contact details           DET  5  FTR 1  low       3
EO  Hall ticket                      DET 14  FTR 3  average   5
EO  Payment receipt                  DET  8  FTR 2  average   5
EO  Registrations by paper report    DET 21  FTR 3  high      7
EQ  View registration status         (rated by the counter)  low       3
EQ  Search the paper catalogue       (rated by the counter)  low       3
unadjusted function points 70; 14 characteristics, TDI 39; VAF 1.04; adjusted function points 72.8
release 2.0's 200 defects: 2.75 per function point

Follow two rows through the rules. The exam registration file has 24 DETs, which OMG's rule places in its middle band (20 to 50), and 3 RETs, also its middle band (2 to 5); the two bands add to 4, which is average, weight 10. Submitting the exam form has 16 DETs, above the external input's middle band of 5 to 15, and 3 FTRs, above its band of exactly 2; the bands add to 6, which is high, weight 6.

munotes.in388

Size Metrics: Lines of Code and Function Points

The thirteen weights add to 70 unadjusted function points. The fourteen characteristics were rated to a total degree of influence of 39, so VAF = 39 × 0.01 + 0.65 = 1.04, and ExamReg is 72.8 adjusted function points. Divided into release 2.0's 200 defects, that is 2.75 defects per function point, a density that can be compared with another system's whatever language each was written in.

Lines of code or function points

Lines of codeFunction points
MeasuresWhat the programmers wroteWhat the software does for its users
AvailableOnly once the code existsFrom the requirements, before design
Depends on the languageYesNo
CountingEasy to automate, once the rule is written downNeeds trained counters and judgement; OMG's standard automates a version of it
Main riskAn unstated counting ruleInconsistent counters; unsuited to some kinds of software

Park weighed function points in 1992 and saw their strengths, "they are language-independent, often solution-independent, and usually computable early in development life cycles, even before specific product designs are available", and three reasons for caution, one of them that "Automated function point counters do not yet exist." That reason has since been overtaken: OMG's Automated Function Points standard is a specification for counting function points from source code by tool. It illustrates this book's rule that a source's claims have a date.

What it does not mean

A line count is not a size until its rule is stated. The same module gave five honest counts from 14 to 30.

Function points do not measure effort or quality. They measure functional size; effort and defects are divided by them.

Adjusted is not always used. The adjustment is a step of the IFPUG method; OMG's automated count and many comparisons use unadjusted function points.

Neither measure is the size. Each measures one attribute; which one to use depends on the question, as the last chapter's goal, question, metric method would ask.

Quick revision

  • Size normalises: defects, effort and cost are divided by it (Park).
  • SLOC (ISO/IEC 20968); Park 1992: physical source lines and logical source statements; a definition checklist states what is included; basic definition excludes comments and blank lines, and a line is classified by its highest-precedence type.
  • Worked example 1: one 30-line module measured 30, 23, 22, 18 and 14 by five rules.
  • Function points: functional size from the user's view (Albrecht, IBM, 1979; ISO/IEC 20926); EI, EO, EQ, ILF, EIF; complexity from DET with RET (files) or FTR (transactions).
  • Weights: ILF 7, 10, 15; EIF 5, 7, 10; EI 3, 4, 6; EO 4, 5, 7; EQ 3, 4, 6.
  • VAF = (TDI × 0.01) + 0.65 over 14 general system characteristics; adjusted = unadjusted × VAF.
  • Worked example 2: ExamReg 70 unadjusted, TDI 39, VAF 1.04, 72.8 adjusted; 2.75 defects per function point.
munotes.in389

Size Metrics: Lines of Code and Function Points

Test yourself

1. Why is "lines of code" an ambiguous measure? How is the ambiguity removed? Because a line count depends on whether blank lines, comments, declarations and continuation lines are counted, and whether physical lines or logical statements are meant; reports of size can be misunderstood by a factor of three or more. It is removed by a written definition, such as the SEI's definition checklist, stating exactly what the count includes and excludes.

2. Distinguish physical source lines from logical source statements. Physical source lines are lines of the source file counted by a stated rule, for example nonblank, noncomment lines. Logical source statements are the language's statements, however many lines each occupies, so one statement written over five lines counts once.

3. Name and define the five components of function point analysis. External input: an elementary process that processes data or control information coming from outside the boundary. External output: one that sends data outside the boundary with processing beyond an inquiry. External inquiry: one that sends data outside the boundary without that extra processing. Internal logical file: a user-recognisable group of related data maintained inside the application. External interface file: such a group referenced by the application but maintained by another.

4. A system has 3 low ILFs, 1 average EIF, 4 average EIs, 2 high EOs and 3 low EQs, and its 14 characteristics total 42. Compute its adjusted function points. Unadjusted: 3 × 7 + 1 × 7 + 4 × 4 + 2 × 7 + 3 × 3 = 67. VAF = 42 × 0.01 + 0.65 = 1.07. Adjusted: 67 × 1.07 = 71.69.

5. Give two advantages of function points over lines of code, and one disadvantage. They do not depend on the programming language, and they can be counted from the requirements before any code exists, so they help estimation. But counting needs trained counters and judgement, and it suits some kinds of software, such as business applications, better than others.

6. What was Park's recommendation on size measures, and why? To start with physical source lines of code, because counters for them are easier and cheaper to build, the rules for logical statements are complex and differ for every language, physical counts are easier to interpret and compare, and most historical data is in physical lines.

munotes.in390

The rest of this subject

These notes are cut from the University's printed syllabus. Open the syllabus itself, or the past papers, for the same subject.

Issue
Done!