Wednesday, October 7, 2026

Investigating the early end of a write on the Virtual 2315 Cartridge Facility

THE ISSUE I AM INVESTIGATING

As I was finishing testing of the Virtual 2315 Cartridge Facility (V2315CF), I was writing a sector of data to verify that my signal quality fixes had resolved earlier problems. The internal 2310 disk drive of the 1130 has sectors of 321 words in size, a word on the 1130 being 16 bits long. 

The disk has 203 cylindrical paths that the heads can access, changed by doing a seek command to move the arm. There is a head on both the upper and the lower surface of the single 14" disk platter inside the removable 2315 disk cartridge. One rotation of the platter is divided into sectors 0 to 3, thus with the arm at a particular cylinder and with one of the heads selected, the system can read or write any of the 4 sectors available. 

I loaded a buffer in the 1130 memory with a block of 321 words that I wanted to write to the disk cartridge loaded onto the V2315CF. I issued a write command on the 1130 and expected that the image of the 2315 disk cartridge held in the RAM of the V2315CF would have the chosen sector from the write command replaced by the words I placed into the memory buffer. 

I then cleared the buffer to all zeroes and issued a read command for the same sector. I expected to see all 321 words exactly matching what I had put in the buffer. However, it matched up to the 159th word out of 321, the remaining words were all zero instead of the values I intended to write.

MODIFIED FPGA LOGIC TO EMIT A DIAGNOSTIC SIGNAL

Among the questions I want to firmly resolve is whether the FPGA logic I designed in the V2315CF is abandoning the write operation at that midway point, or whether the issue is coming from the IBM 1130 itself. I set up a diagnostic signal that will be 1 when the write state machine is active and routed this out on the external signal lines on the V2315CF so that I could see it on an oscilloscope or logic analyzer.

WATCHING KEY SIGNALS ON THE 1130 PLUS THE DIAGNOSTIC SIGNAL

I put the oscilloscope on the control signal -Write Gate produced by the 1130 when it is writing to the sector. I also hooked up to the -Sector Marker signal to see where the failure occurs in relationship to the sector marking pulses. I recorded the combined output signal -Write Clock and Data from the 1130 to see if it is emitting the intended words after number 159 or is sending words of all zero. 

The -Write Gate signal stays high for the entire 10 milliseconds duration and the sector mark signal seemed reasonable. I wired up the logic analyzer to twelve signals and recorded it to help figure out what is happening here. The signals were:

  • -Write Gate
  • -Sector Marker
  • -Four Sector Pulses
  • -Int Req L2
  • -Full Wd Count
  • -Req CS 0
  • -Index Marker
  • -Increment Word Count
  • -Write Clock Phase B
  • -Write Clock and Data
  • -RW Condition
  • +Sector Marker

TIMING OF KEY SIGNALS CONFIRMED FROM LOGIC ANALYZER

One of the concerns I had was whether the various single shot timers were set properly to allow adequate time for the completion of a write of a sector. Fortunately, the timing is very good, thus no need to readjust the timers.

The sector marker starts the sector. The marker should go low for 160 microseconds, which is generated by my logic in the FPGA of the main unit of the V2315CF. A sector on the disk consists of one fourth of a revolution of the disk platter, marked with two sector marker pulses 5 milliseconds apart. Every other sector marker pulse is blocked, producing the signal -Four Sector Pulses which activates once every 10 milliseconds to delineate the next sector.

When the sector marker pulse ends, 160 ms into the sector, a timer of about 250 microseconds allows the clock and data pulses to transmit a long string of zero bits called the preamble. At the end of the preamble, the signal -RW Cond turns on for the remainder of the write activity. This happened at almost exactly the intended 250 microsecond from the end of the -Sector Marker pulse. 

The end of the preamble comes after we transmit the sync word, a data word of 0x8000 which is recorded on the disk with the last bit first. A word on the 1130 is 16 bits long, which is followed by four error checking bits that enforce the rule that the number of bits that are 1 must be an even multiple of four. Thus the ECC bits can be 0000, 1000, 1100 or 1110 depending on how many 1 bits are needed to reach the even multiple. The sync word needs three more 1 bits after the data is transmitted, so the sync word appears as a 1 followed by the ECC pattern 1110. 

Immediately after, data words are pumped out, with each 20 bits representing a data word plus the ECC bits. The sector consists of 321 words. After each word is completely transmitted, the signal -Incr Wd Count will blip low to count down by 1. When the count goes down from 321 to zero, meaning we have finished sending out the data, the signal -Full Wd Count will go low and end the sector. That also drops the -Int Req L2 signal to interrupt the CPU as the write command has completed. 

As each word is being transmitted, the signal -Req CS 0 will go low to ask for the next word to be fetched from the memory of the 1130. Everything in the trace above looks correctly and properly timed. We also see actual data patterns in the -Write Clock and Data signal for the words being written.

SEEING WHERE THE LIVE DATA STOPS BEING WRITTEN

Looking at the end of the trace, the last word is requested from memory, then the word count is decremented to zero, the interrupt request level is asserted, and the -Write Gate is turned off. This represents the correct end of the write operation, however I can see that the data words are all zeroes for about half of the sector. 


In the trace above, you see the request for the last word, the increment of the word counter and then the -Full Wd Count signal drop. That turns off -RW Cond and turns on -Int Req L2 as the write operation is done. The next -Sector Marker starts which is the end of our sector and causes -Write Gate to return high. All looks correct except that our data has been zeroes rather than the values in the 1130 memory. 

I looked at the point where I stopped seeing valid data words being emitted. It is soon after the -Sector Marker pulse, but that one is not passing through as -Four Sector Pulses so it should not be affecting the disk controller logic. 


I circled the -Req CS 0 signal that fetches the first word that is erroneously zero. I had loaded the memory buffer with 321 words each having its word number as the data value, thus none should be zero. The green circle shows the bits being shifted out have a mix of 1 and 0 data values - these are legitimate. However, once the ECC bits of the last correct word are shifted out, we see the pattern of nothing but zeros for all the remaining data words. 

The duration of the sync word plus 321 data words is 9.347 milliseconds. The first bad word takes place at about 4.6 ms into that duration which is just about the halfway point of the 321 words. That doesn't correspond to an even address multiple, nor even binary count value that would correspond to a bad signal line or gate that would cause us to mis-address memory. 

I decided to decode the last few words that were output and I believe I have a clue. Starting with the last non-zero word and going backwards, I saw x013C, x013A, x0138, and decreasing by twos. That was the clue! My address is jumping by two rather than one, which is why the data ran out early. What is funny is that the early data words step one address at a time, so something has happened to increase the stride from 1 to 2. 

The logic in the 1130 increments the address at the completion of fetching a word, when the signal +CS Level 0 drops. The dance that occurs fetching the data begins when the prior word has had its 16 data bits transmitted over -Write Clock and Data to the disk drive, but prior to the four ECC bits being transmitted. A request to fetch a word from memory (a cycle steal) is asserted by dropping the -CS Req 0 that asks for the top priority memory fetch. 

Once the cycle steal begins, signal +CS Level 0 goes high to report that. During that 3.6 microseconds of a memory cycle, at about 1.3 microseconds from the start, the data has arrived and is being stored in the register that shifts the data bits out to the disk drive. We have four ECC bits to shift out, requiring 11 microseconds which allows plenty of time for the cycle steal to overlap it and the new data to be ready in the register. 

When the cycle steal ends, the address register should increase by 1. A bit later the signal -Increment Word Count goes low to cause the count of words transmitted to be reduced, eventually the count reaches zero and the -Full Word Count goes low to end the write operation. 

If there is a hardware defect in the 1130, it is likely in the circuitry that holds the memory address (the File Address Register or FAR) inside the disk controller logic of the 1130. That circuit also performs the addition when the cycle steal signal +CS Level 0 drops back to low at the end of the fetch. 

The register is a chain of flipflops with combinatorial gates to set and reset each flipflop. There FAR can be reset with the -Reset File Address signal, loaded from memory at the start of the write operation with the -Load File Address signal, and incremented when +CS Level 0 drops low at the end of each memory fetch. The maximum size of an 1130 system is 32K words, thus we need 15 address bits for the FAR; on the 8K system I am restoring, only 13 bits are necessary but the design encompasses the largest configuration. 

In order to understand the circuit, the peculiarities of the IBM logic family used in the 1130 and the documentation methodology they use has to be explained first. The flipflops used in the 1130 are set or reset by the falling edge of an input, not by a steady state signal level. 


The outputs are both a regular and an inverted output, shown at the top and bottom of the right edge respectively. The outputs are steady state other than for undesirable pulses that occur if a flipflop is already set and receives another set input pulse, as well as when the flipflop is unset and receives another reset input pulse. The flipflop produces a brief pulse on the opposite output line - e.g. resetting a flipflop that is already in the off state will produce a brief spurious pulse on the positive output line. 

IBM designs around the spurious output pulses by using combinatorial logic gates that produce or block the input pulses that would cause this issue. The combinatorial gate used to create the set or reset pulses for the flipflop are what IBM called AC triggered gates. 

In the image above, there are two AC triggered gates which will produce a brief negative going pulse when they activate. These are combined together to produce the final pulse that is used to set a flipflop, so that if either of the gates on the left fire, its output sets the flipflop. An AC triggered gate will activate at the falling edge of one input trigger signal, marked with an N on the left side of the gate, but only if the conditioning (other) input to the gate is steady low at the time the trigger pulse. 

Thus, if the conditioning input -B Bit 15 is low at the time that -Load File Address has a falling edge, the lower gate will produce a set pulse for the flipflop. Similarly, if the flipflop is not set, then +File Address Bit 15, the output of the flipflop, is low and conditions the upper gate. When +CS Level 0 has its falling edge, the flipflop is set but only if it had been unset before. 


Thus our flipflop for FAR bit 15 will be set if it is not previously set, when the cycle steal memory fetch ends, or will be set if the value on B Bit 15 is low when we are loading the file address. -B Bit 15 is inverted, so it is low when bit 15 contains a 1. The lowest gate on the left simply carries the signal -Reset File Address through to the reset input of the flipflop; a negative pulse on the reset file address line will reset the flipflop immediately. 

This means we set the FAR flipflops to match what is in the B register (memory fetch output) when we drop the signal to load the file address. The flipflop will be reset when the cycle stead memory fetch ends but only if it had previously been set. We also can reset this if the signal -Reset File Address is dropped low. 

The effect of the +CS Level 0 signal is to toggle the state of the flipflop each time that we have a falling edge on the input signal - every time we complete a cycle steal memory fetch. This is the first step in the addition process that increments the FAR after each cycle steal.

Whenever a bit of the FAR goes off, it toggles the next higher bit, which is how we make this increment. Below, you see that the output of +File Address 15 is connected to the trigger of the gates for FAR bit 14, so that whenever bit 15 turns off, the falling edge will toggle the state of bit 14. The remaining bits are similarly connected, thus dividing by two on each bit position and yielding an address that steps by one every time we fetch a word from memory via cycle steal. 


FAR bit 14 has the same logic to load or reset in addition to its role toggling its state with each drop of bit 15 to zero. This all seems pretty straightforward and the logic analyzer traces show only a single falling edge of +CS Level 0 occuring on each word. Based on this, unless something is defective in the circuits, this should count up strictly by a stride of 1. Further, nothing I see in the circuit would make the low bit (FAR 15) jump based on the state of the upper bits. 

My only speculation is the part of the setting logic for FAR 15 that looks at the value in B bit 15 and turns on the flipflop when the -Load File Address signal drops. If this is failing, then any data word fetched from memory with bit 15 having a 0 value might set the flipflop, then the end of the cycle steal fetch would toggle it a second time. A mechanism like this would produce a stride of 2, which is supported by the data values of the final few words fetched which all had the low order bit value of 0. 

NEW LOGIC ANALYZER RUN TO LOOK FOR THE CAUSE

Since it appears that something goes wrong in the circuits that increment the FAR, I will capture those with the logic analyzer and look for a) confirmation that it double steps FAR bit 15 or fails to bump FAR bit 15, as well as b) indications of the signals that might contribute to that failure. Hopefully that will let me dive into the issue and look for the true cause. 


No comments:

Post a Comment