Sharing

顯示具有 Storage 標籤的文章。 顯示所有文章
顯示具有 Storage 標籤的文章。 顯示所有文章

2012年2月3日 星期五

Evolution of the Storage Brain 筆記 (二)


Chapter 4. A Journey to the Center of the Storage Brain

The Past: Refrigerators and Cards
  • Skip
The Middle Ages: The Age of RAID Controllers

    RAID controllers combined the basic functionality of disk controllers with the ability to group drives together for added performance and reliability
    • Communication between the underlying disks and the attached host computer
    • Protection against disk failure through the creation of one or more RAID groups
    • Dual-controller systems, which could be used for added performance, availability and automated failover
    The Modern Age: Storage Array Controllers
    • The storage industry became very volatile during the 90s
    • Modern-day networked storage was born. 
    • EMC and NetApp capitalized on increased storage intelligence and grew at record paces
    Enterprise Array Controllers: A More Intelligent Brain

    As the storage industry moved from the late 90s into the mid 00s, the market stabilized into a smaller albeit more mature set of suppliers

    Table: Data Management Tasks Addressed by Modern Arrays

    Function
    Description
    Performance
    Storage arrays are required to quickly process a multitude of data I/O requests arriving simultaneously from hundreds (or thousands) of desktops and servers
    Resiliency
    Ø  Automated RAID Rebuilds
    Ø  Data Integrity Checks
    Ø  “Phone Home” Alerts
    Ø  Environmental Monitoring
    Ø  Non-disruptive Software Updates
    Virtualization
    Ø  Transparent Volume/LUN resizing
    Ø  Thin Provisioning
    Ø  Data Cloning
    Ø  Data Compression
    Ø  Data Deduplication

    Chapter 5.  Without Memory, You Don’t Have a Brain

    Human
    PC
    Sensory memory
    buffers and registers
    Short-term memory
    cache, Random Access Memory (RAM), flash and solid-state disks (SSDs)
    Long-term memory
    hard disk drives

    Era
    Memory Type
    40s-50s
    Cathode Ray Tubes
    50s-60s
    Magnetic Core Memory
    70s-present
    Random Access Memory (RAM) integrated circuits

    NetApp add some interesting algorithm-level intelligence to its PAM cache
    • Priority-based caching
    • Non-redundant data caching
    • Predictive caching
    • Immediate caching
    • Metadata caching
    Here are the trends we’ll see in the industry’s efforts to reach this goal
    • Application servers and workstations will cache more and more of their own data
    • Storage networking switches will cache more and more data as it travels through their path
    • Storage systems will cache more and more front-end data
    • SSDs will become the first line of defense for large amounts of data that can’t be stored in cache
    • Hybrid disk drives will be the final step
    Chapter 6. The Storage Nervous System

    Different between NAS, SAN, and iSCSI, 看了這麼多解釋, 我還是覺得鳥哥的圖最棒




    NAS Communications Protocols
    NFS
    NFS stands for the Network File System protocol used by UNIX-
    based clients or servers
    CIFS
    CIFS stands for Common Internet File System. CIFS is another
    network-based protocol used for file access communications
    between Microsoft Windows clients and servers.
    SAN Communications Protocols
    Fibre Channel Protocol (FCP)
    Typically occurs via specialized Fibre Channel cabling, Fibre Channel host bus adapters (HBAs) and Fibre Channel switches operating between the SAN and its hosts
    iSCSI
    iSCSI stands for Internet Small Computer Systems Interface. iSCSI is “a transport protocol that provides for the SCSI protocol to
    be carried over a TCP-based IP network.”
    FcoE
    FcoE stands for Fibre Channel over Ethernet. FcoE allows Fibre Channel storage traffic to be sent over Ethernet networks

      What is data package collision in network? 

      • A network collision occurs when more than one device attempts to send a packet on a network segment at the same time
      • Collisions are resolved using carrier sense multiple access with collision detection in which the competing packets are discarded and re-sent one at a time
        • This becomes a source of inefficiency in the network
      • Collision domains are found in a hub environment
      • Collision domains are also found in wireless networks such asWi-Fi.
      • Modern wired networks use a network switch to eliminate collisions.

      Early Heritage Evolved into Separate Paths & Growing Confusion


      10 Megabit per second (Mbps) 
      Ethernet was the norm with 100 Mbps Ethernet just 
      emerging. 
      SANs came out of the gate with fiber optic cables and a protocol that could move data at 1 Gigabit per second (Gbps), a ten-fold improvement over the fastest Ethernet-based NAS transport
      • anyone with a “need for speed” simply had to use SAN storage
        • Databases, transaction processing, analytics, and similar applications fell into this category
      • NAS, although less costly and easier to implement than SAN, was usually relegated to “slower” applications
        • User files, images and Web content
      New Realities Point to SAN/NAS Convergence and Unification 


      10-Gigabit Ethernet (10-GbE or 10 Gbps) is common. 100-Gigabit Ethernet (100-GbE or 100 Gbps) devices have also been demonstrated.  Fibre Channel networks operating at 4-Gbps are common today, with 8-Gbps having also been newly delivered.


      • Unifying block and file data transport onto the same network fabric
      • Unifying SAN and NAS data storage functionality onto the same storage system.



      Advancing Intelligence for Internal Communications

      SAN Virtual Storage
      • add a second logical abstraction layer to the physical disk drives
      • did not map LUNs to physical drives, it mapped LUNs to the logical blocks stored on these drives
      NAS Virtual Storage
      • NetApp had created an intelligent internal communications network specifically designed for storage systems. The WAFL file system provided a communications network based on logical file-based objects



      Next breakthrough in storage protocols and communication becomes object-based storage, SONET, RDMA






























      2012年2月1日 星期三

      Evolution of the Storage Brain 筆記 (一)

      Chapter 1. And Then There Was Disk


      First came to market in the late 50s and 60s, disk drives have relied on the following core components

      • Read/write heads that use electrical impulses to store and retrieve magnetically recorded bits of data
      • Magnetically coated disk platters that spin and house these bits 
      • Mechanical actuator arms that move the heads back and forth across the spinning disk platters, forming concentric ‘tracks’ of recorded data
      The Past: Disk Drives Prior to 1985

      The Early Days of Disk Drive Communications
      • Bus-and-Tag Systems
        • via two copper-wire cables: Bus and Tag, Bus for data, Tag for communication protocols
      • SMD Disk Drives
        • Control Data Corporation first shipped its minicomputer with storage module device (SMD) disk drives in late 1973
        • Used much smaller ―A and B flat cables to transfer control instructions (from the A cable) and data (from the B cable)
        • The disk controller shrank down to a single board which was inserted into the system‘s CPU card cage
      The Middle Ages: Disk Drives in the 80s -90s


      Disk drives in the late 80s and 90s went through a number of significant transformations that allowed them to be widely used in the emerging open systems world of servers and personal computers. These included:

      • 19” disk drives => less expensive 5.25” (and, eventually, 3.5”) drives.
      • Advances associated with redundant arrays of independent disks (RAID) technology. 
      • The development and widespread adoption of the Small Computer Systems Interface (SCSI). 

      Disk drives produced today fall into four categories, depending on their cable connections
      • SCSI
      • Fibre Channel
      • Serial ATA (SATA)
      • Serial Attached SCSI (SAS)
      SCSI used a single data cable to present its Common Command Set (CCS) interface with built-in intelligence.


      Today: Disk Drives, Pork Bellies and Price Tags

      Are Disk Drives in Our Future?

      Today‘s latest battle cry is that solid state disks (SSDs) will completely replace magnetic disk storage

      Research into the area of higher capacities for magnetic disk drives

      • Perpendicular Magnetic Recording (PMR)
        • 多層次的儲存,原本是平面,變成是 3D
      • Patterned Media Recording
      • Heat-Activated Magnetic Recording
        • relies on first heating the media so that it can store smaller bits of data per square inch
        • PMR appears to be winning the short-term race
      • Nanostorage
      Chapter 2. “Oh, @#$%!” 

      Address protection in two separate chapters
      • Protecting against disk drive failure
        • users can continue to access data previously stored on failed disks
      • Protecting against data loss or corruption
        • moves more deeply into storage software intelligence

      The Past: Protecting SLEDs

      When a drive crashed, data was recovered from tape and restored back to a new disk


      The Middle Ages: RAID in the 80s

      1987 paper 
    • called, “A Case for Redundant Arrays of Inexpensive 
    • Disks.” RAID was born.

      • provide greater efficiency and faster I/O performance
      • successfully survive a failure of any one disk drive
      • describe five different RAID methods (RAID 1 through RAID 5)


      RAID Technique
      Description
      No Parity
      (RAID 0)
      Ø  increase I/O performance by striping  (or logically distributing) data across several disk drives.
      Ø  offered no protection against failed disk drives
      Mirroring
      (RAID 1)
      Ø  data is mirrored onto a second set of disks
      Ø  exacts a high capacity penalty
      Fixed Parity
      (RAID 3 / RAID 4)
      Ø  both use parity calculations (sometimes known as checksum) to perform error-checking
      Ø  recovery of missing data from failed drives
      Striped Parity
      (RAID 5)
      Ø  parity is striped (or logically distributed) across all disks in the RAID set in an attempt to boost RAID read/write performance
      Multiple Parity
      (RAID 6)
      Ø  using multiple iterations of fixed or striped parity on a group of drives, which allows for multiple drive failures without data loss.


      The Future: Smarter, Self-Healing Disk Drives

      S.M.A.R.T. technology

      Chapter 3. Virus? What Virus? 

      Approaches to Data Loss or Corruption


      Approach
      Description
      Data replication
      Mirroring critical data to an alternate location
      Data backup
      Restore data that may have been accidentally deleted or earlier data version


      The Past: The Tale of the Tape


      Today’s Backup Tapes


      Decades of “format wars” ensued amongst vendors fighting for market share. Sample formats from this era included:



      • Quarter-Inch Cartridge (QIC)
      • 4mm or 8mm Tape
      • Digital Linear Tape (DLT)
      • Advanced Intelligent Tape (AIT)
      • Linear Tape Open (LTO)


      LTO-4 has become the reigning tape format today.

      The Middle Ages: Early Tape Backup Automation


      Tape-Based  Innovations:
      Interleaving
      Ø  improved tape backup speeds by allowing backups to be written to multiple tapes concurrently
      Ø  writing to several tapes in parallel
      Synthetic Backup
      Ø  required just a single full backup and used an intermediate database to track and map the location of the continuous incremental backups performed to tape thereafter
      Reclamation
      Ø  Also pioneered by Tivoli Storage Manager (TSM)
      Ø  the tape reclamation process solved a problem created by Synthetic tape backups
      Disk-Based Innovations
      Disk Staging
      Ø  data stored on optical media that could be moved by a staging manager to magnetic disk drives
      Ø  could be used to send the data to tape without affecting production workloads
      D2D2T
      disk-to-disk-to-tape
      D2D
      disk-to-disk without tape


      The Modern Age: Emerging D2D Efficiencies

       A Snapshot is Not a Backup


      NetApp explains this space-saving functionality as follows
        • We are able to create a snapshot in constant time because we have a map of the blocks that are allocated on disk. A snapshot is really just a copy of the block map rather than the actual disk blocks



        Use of local snapshots alone, however, still exposes the data to other corruption risks, such as

        • Widespread data corruption of the primary data set
        • Hardware failure impacting the data stored within
        How SnapVault Works:
        • The SnapVault “primary” system needing data protection.
        • The SnapVault “secondary” system where backup data is stored.
        1. SnapVault initially stores one “full” backup of the primary system‟s data set on the secondary
        2. builds on NetApp Snapshot efficiencies by quickly transmitting only the changed blocks found in the most recent snapshot of the primary system


        The Future of Data Protection


        • The new gold standard: Annual off-site archival of data to tape
        • Tape backup will become a service in lieu of local tape libraries.
        • D2D will be managed by the storage array itself.
        • Say goodbye to VTLs







        2012年1月4日 星期三